OpenAI prepares new $500/month Pro Max plan for ChatGPT (2 minute read)
OpenAI appears to be preparing a new ChatGPT Pro Max subscription priced at $500 per month. It is unknown whether the company plans to justify the $500 price through speed, larger usage allowances, longer-running Work sessions, or a combination of the three. The timing of the leak is notable as OpenAI's DevDay will take place on September 29. The company is expected to discuss APIs, developer tooling, new subscription tiers, and new products at the event.
|
Bringing Your Muse to Life (5 minute read)
Meta has introduced Muse Realtime Avatar, a technology that transforms conversations into expressive, interactive avatars in real time. Using a seamless integration of voice and visual performance, it synchronizes speech and avatar expressions accurately in live interactions. Muse Realtime Avatar outperforms competitors with improved visual quality and real-time responsiveness, supported by Meta's optimized AI inference stack.
|
Gemini 3.8 Live with Live Avatar (1 minute read)
Google launched Gemini 3.8 Live featuring a Live Avatar, enhancing user interaction. This update offers real-time responses and more personalized user experiences. It positions Google to compete more aggressively with other AI platforms.
|
|
Modern LLMs have tiny GPTs hidden inside them (8 minute read)
It's likely that modern LLMs have tiny self-models of LLMs inside of them that allow them to better predict the next token generated by LLMs. Models approximating LLMs within themselves would open up all sorts of metacognitive abilities. It would open doors for a self-model within LLMs similar to a model humans have about themselves they call a 'self'.
|
700 TPS on Kimi K3: A Case for TPU Megakernels (20 minute read)
inferact/tpu-megakernels is a collection of megakernels for TPU v7. Its Kimi K3 implementation delivers over 700 tokens per second with speculative decoding. Without speculative decoding, its megakernels for K3 and Qwen 3.8 27B deliver roughly 1.4 to 2x the decode throughput of the GB200 baseline at batch sizes 1 through 8. This post explains the design and why megakernels and TPUs are a good fit.
|
Why the Best AI Clouds Don't Run on Flash Alone (8 minute read)
Flash storage is fast and reliable, but much of the storage work that surrounds active training doesn't need flash-grade speed or prices. A lot of the infrastructure neoclouds have is built around flash, so for storage that doesn't need flash, they often point customers to hyperscaler object storage solutions instead. The more data you have stored somewhere, the harder it becomes to leave. While the neocloud keeps getting paid for GPU, the hyperscaler quietly takes over the broader account.
|
|
Contrastive Language Models (8 minute read)
Contrastive Language Models (CLMs) are a new class of System One model trained with a contrastive learning objective that connects states and actions. CLM-8B delivers performance comparable to Jev across computer-use, gaming, and tool-calling tasks while achieving up to 9x lower latency. It also sets a new state-of-the-art on changing agentic coding benchmarks. The model is pre-trained on 60M Nemotron Q&A pairs, mid-trained on 30M synthetic hard negatives, and post-trained on 1M agentic trajectories.
|
Managed Deep Agents v0.8: new auth, memory, and channels (12 minute read)
Managed Deep Agents combines the Deep Agents harness with all of the infrastructure required to run agents in production. It is the simplest way to build, deploy, and run mission-critical agents in production. Managed Deep Agents 0.8 adds support for user-owned credentials, user-level memory, HTTP channels, and file transfer in Slack. It also adds a pre-built tool for web search powered by Parallel.
|
|
Are you ready for superintelligence (13 minute read)
Recent frontier models are rapidly saturating old benchmarks and moving into harder real-world, scientific, and agentic tasks. The pace of capability gains, alongside soaring AI usage and revenue, is shifting debate from workplace augmentation toward recursive self-improvement and superintelligence.
|
|
|
|
|