TLDR AI
{{PreviewText}} ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌  ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ 

TLDR

Together With AMD

TLDR AI 2026-09-25

Take control of your AI costs with AMD (Sponsor)

As AI agents take on more work, cloud inference costs can grow with every task. The AMD Tokenomics Calculator helps you understand the economics of your AI workloads and optimize token spend.

  • Understand your AI costs. Model infrastructure costs to see how token usage translates into ongoing spend.
  • Unlock the economics of local AI. See how running AI locally can reduce cloud inference costs and decrease token spend.
  • Reduce recurring AI costs. High-performance agentic PCs featuring AMD processors run multiple agents locally, completing tasks up to 6x faster while reducing reliance on cloud inference.

👉 Calculate your potential savings now

🚀

Headlines & Launches

OpenAI prepares new $500/month Pro Max plan for ChatGPT (2 minute read)

OpenAI appears to be preparing a new ChatGPT Pro Max subscription priced at $500 per month. It is unknown whether the company plans to justify the $500 price through speed, larger usage allowances, longer-running Work sessions, or a combination of the three. The timing of the leak is notable as OpenAI's DevDay will take place on September 29. The company is expected to discuss APIs, developer tooling, new subscription tiers, and new products at the event.
Bringing Your Muse to Life (5 minute read)

Meta has introduced Muse Realtime Avatar, a technology that transforms conversations into expressive, interactive avatars in real time. Using a seamless integration of voice and visual performance, it synchronizes speech and avatar expressions accurately in live interactions. Muse Realtime Avatar outperforms competitors with improved visual quality and real-time responsiveness, supported by Meta's optimized AI inference stack.
Gemini 3.8 Live with Live Avatar (1 minute read)

Google launched Gemini 3.8 Live featuring a Live Avatar, enhancing user interaction. This update offers real-time responses and more personalized user experiences. It positions Google to compete more aggressively with other AI platforms.
DeepSeek allegedly doubled its revenue run rate to $1 billion (1 minute read)

DeepSeek allegedly doubled its annualized revenue run rate to $1 billion after raising API prices by 2.3 to 4.5 times. The Information says developer demand held up despite the increase.
🧠

Deep Dives & Analysis

Modern LLMs have tiny GPTs hidden inside them (8 minute read)

It's likely that modern LLMs have tiny self-models of LLMs inside of them that allow them to better predict the next token generated by LLMs. Models approximating LLMs within themselves would open up all sorts of metacognitive abilities. It would open doors for a self-model within LLMs similar to a model humans have about themselves they call a 'self'.
AI may learn to reason in representations people cannot read (20 minute read)

Models may reason through internal states that never appear in their written chain of thought, then give a plausible explanation after the fact. That would weaken safety checks that rely on reading the visible reasoning.
700 TPS on Kimi K3: A Case for TPU Megakernels (20 minute read)

inferact/tpu-megakernels is a collection of megakernels for TPU v7. Its Kimi K3 implementation delivers over 700 tokens per second with speculative decoding. Without speculative decoding, its megakernels for K3 and Qwen 3.8 27B deliver roughly 1.4 to 2x the decode throughput of the GB200 baseline at batch sizes 1 through 8. This post explains the design and why megakernels and TPUs are a good fit.
Why the Best AI Clouds Don't Run on Flash Alone (8 minute read)

Flash storage is fast and reliable, but much of the storage work that surrounds active training doesn't need flash-grade speed or prices. A lot of the infrastructure neoclouds have is built around flash, so for storage that doesn't need flash, they often point customers to hyperscaler object storage solutions instead. The more data you have stored somewhere, the harder it becomes to leave. While the neocloud keeps getting paid for GPU, the hyperscaler quietly takes over the broader account.
🧑‍💻

Engineering & Research

Are businesses closing the AI observability gap? (Sponsor)

1 in 4 AI agents still deploy without runtime monitoring, per New Relic's Observability Forecast survey of 2.5k IT and engineering pros. And as organizations scramble to close the gap, tool sprawl is creeping back in. To see what sets leaders apart from those playing catch-up, browse the report
Contrastive Language Models (8 minute read)

Contrastive Language Models (CLMs) are a new class of System One model trained with a contrastive learning objective that connects states and actions. CLM-8B delivers performance comparable to Jev across computer-use, gaming, and tool-calling tasks while achieving up to 9x lower latency. It also sets a new state-of-the-art on changing agentic coding benchmarks. The model is pre-trained on 60M Nemotron Q&A pairs, mid-trained on 30M synthetic hard negatives, and post-trained on 1M agentic trajectories.
Managed Deep Agents v0.8: new auth, memory, and channels (12 minute read)

Managed Deep Agents combines the Deep Agents harness with all of the infrastructure required to run agents in production. It is the simplest way to build, deploy, and run mission-critical agents in production. Managed Deep Agents 0.8 adds support for user-owned credentials, user-level memory, HTTP channels, and file transfer in Slack. It also adds a pre-built tool for web search powered by Parallel.
🎁

Miscellaneous

Product Manager, Applied AI at TLDR ($200k base + $60k bonus, Fully Remote)

TLDR is hiring its first PM to help build the agent-first operating layer used across the company. We're looking for a builder who has shipped real products/systems with LLMs. Click here to learn more.
What we learned from being the first company to disclose an agent cyberattack (4 minute read)

Hugging Face says its autonomous-agent cyberattack exposed three priorities: stronger incident transparency, reducing capability asymmetries between attackers and defenders, and preserving open-source access for defense. AI created new attack risks but also helped investigate, mitigate, and harden systems.
Are you ready for superintelligence (13 minute read)

Recent frontier models are rapidly saturating old benchmarks and moving into harder real-world, scientific, and agentic tasks. The pace of capability gains, alongside soaring AI usage and revenue, is shifting debate from workplace augmentation toward recursive self-improvement and superintelligence.
⚡

Quick Links

Anthropic tested what happens when agents bargain for people (38 minute read)

After five-minute interviews, Claude agents traded books for employees, and their preference rankings matched the humans' on 61% of pairs.
How Good Are LLMs at Decision Forking? (GitHub Repo)

Taste-Bench evaluates whether an LLM agent can choose the better next step at consequential forks in long-horizon tasks.
Intelligence Density (5 minute read)

Trajectory.ai focuses on "intelligence density," measuring cost per task instead of cost per token for more efficient AI model training.
A finance benchmark asks agents to finish the whole assignment (18 minute read)

DAYJOB: Finance tests whether AI agents can complete 80 realistic finance assignments using supporting documents and produce analysis a professional can use.

Want to advertise in TLDR? 📰

If your company is interested in reaching an audience of AI professionals and decision makers, you may want to advertise with us.

Want to work at TLDR? 💼

Apply here, create your own role or send a friend's resume to [email protected] and get $1k if we hire them! TLDR is one of Inc.'s Best Bootstrapped businesses of 2025.

If you have any comments or feedback, just respond to this email!

Thanks for reading,
Andrew Tan, Ali Aminian, & Jacob Turner


Manage your subscriptions to our other newsletters on tech, startups, and programming. Or if TLDR AI isn't for you, please unsubscribe.