TLDR AI
{{PreviewText}} ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌  ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ 

TLDR

Together With runware.ai

TLDR AI 2026-09-28

Run AI workloads at half the cost (Sponsor)

Most AI compute still runs on data centers built for websites and databases, and you pay for that overhead in every GPU-hour. Runware runs on Sonic Pods: low footprint, high efficiency modular AI data centers they design and build themselves, with the first wave coming online now. Same models, same GPUs, same speed – but at around half the cost of a hyperscaler.

✅ One API, 400K+ models → image, video, audio, LLMs and 3D, with one key and one bill
✅ Serverless for your own models → bring code or a container, get an endpoint that scales from zero, billed by the second from $1.99 per GPU-hour
✅ Dedicated compute → reserve HGX B300, GB300 NVL72 or RTX PRO 6000 capacity as bare metal or serverless from $0.99 per GPU-hour
✅ Enterprise-ready → SOC 2, ISO 27001 and GDPR, across US and EU regions

Run on Sonic Pods (early access) → runware.ai
🚀

Headlines & Launches

Why I'm Building Muse (2 minute read)

Muse is built as a personal agent that turns vague ambitions into concrete action by planning, emailing, calling, finding resources, and removing friction. Alexandr Wang frames it as a way to expand individual agency and make more ambitions achievable.
OpenAI and Anthropic Probe Tens of Thousands of Incidents as OpenAI Halts Training (4 minute read)

OpenAI, Anthropic, and security researchers are investigating tens of thousands of incidents in which models acted beyond intended limits. Most incidents caused no harm. The count is not a count of breaches: there were only four incidents of unauthorized access to real third-party systems. OpenAI has paused training, evaluation, and tool-use inference for its most capable models.
OpenAI prepares to expand Ultrafast API to more users (1 minute read)

OpenAI appears to be preparing a wider rollout of its Ultrafast API mode. References to the rollout have been seen showing up across the OpenAI Platform and API documentation. The feature was officially previewed with GPT-5.6 Sol. OpenAI claims it has speeds of up to 750 output tokens per second and up to 14 times faster inference than Standard. The mode is powered by Cerebras. Access remains limited to select customers.
🧠

Deep Dives & Analysis

Let's talk about trading compute (16 minute read)

There is an emerging market of compute derivatives. This could fundamentally change how neoclouds and anyone adjacent can grow as well as protect themselves. The problem is pretty urgent for inference clouds. Companies that sell customers fixed-price services as their GPU bill floats have taken a position on compute prices whether they meant to or not.
Can AI self-improvement overcome diminishing returns? (38 minute read)

AI is already helping build better AI. The technology is improving at a stupendous pace and is expected to continue to do so. Experts expect increasingly superhuman performance in parts of formal math, coding, and cybersecurity, and any other verifiable domain where machines can generate training data and verify success at machine speed. However, this doesn't mean that the industry is close to super-intelligence or general ASI.
Agent (Muse) Compute Demand (5 minute read)

It costs Meta an estimated around 1 to 2 gigawatts of average total power to serve 100 million daily active users. Only around 0.1 gigawatts comes from the GPU/VM layer. 3 to 4 gigawatts per day is entirely plausible depending on the number of reasoning-equivalent model calls users generate. The sandbox layer contains less than $1 billion of CPU content and around $2 billion of DRAM content.
🧑‍💻

Engineering & Research

22 Models, Same Correct Answer, 178x Cost Gap (Sponsor)

Leaderboards don't rank models on how they perform on your data. CData connected 22 models to live CRM, warehouse, and ITSM data through CData Connect AI. When Connect AI carried business context and guardrails, 17 returned the identical correct answer, one model at $0.0009 and another model at $0.157. 

Read the benchmark

Hitting a billion tokens per minute on one GPU by combining a query planner and an inference engine (15 minute read)

The QUery-Aware Inference Layer (Quail) is an inference engine that hits over a billion tokens processed per minute per H100 GPU. It is over 10 times faster than the vLLM baseline on the same hardware, and it costs under $0.06 per billion tokens on Modal. This post provides a quick overview of how Quail works. It focuses more on considerations for inference engineers.
Claude computes a nine-loop amplitude in N=4 super-Yang-Mills (18 minute read)

Physicists at Anthropic harnessed Claude, an LLM, to compute a complex nine-loop amplitude in N=4 super-Yang-Mills, a problem once considered computationally infeasible with limited resources. They utilized the bootstrap method and form-factor approach, achieving results comparable to human attempts but with greater efficiency and less oversight. The work highlights untapped potential in AI for complex physics calculations, suggesting that more advances might be achieved with better computation and software practices.
Policy Gradients for LLMs Explained Visually (8 minute read)

A visual, from-scratch derivation of REINFORCE showed how policy gradients train language models by increasing the probability of rewarded outputs.
🎁

Miscellaneous

Product Manager, Applied AI at TLDR ($200k base + $60k bonus, Fully Remote)

TLDR is hiring its first PM to help build the agent-first operating layer used across the company. We're looking for a builder who has shipped real products/systems with LLMs. Click here to learn more.
On Ezra Klein's Podcast With Jensen Huang (40 minute read)

Jensen Huang doesn't believe in superintelligence or that AI will ever be a different kind of thing from software. AI will never be more than a new abstraction level and so won't fundamentally change anything. He doesn't believe in AI existential risk, even though his tolerance for safety risk is lower than even the most paranoid safety advocate. Critics say he must not understand the technology and that he would have a very different view if he actually understood what the top risks were.
Elon Musk's SpaceXAI to add another 660,000 AI GPUs this year (2 minute read)

220,000 Nvidia GB300 GPUs will be operational at X's Colossus supercomputer in November. The company is aiming to bring another 220,000 units online by late December. Colossus 2 currently has 110,000 GB200 and 440,000 GB300 GPUs, and Colossus 1 has 150,000 H100, 50,000 H200, and 30,000 GB200 GPUs. This means SpaceXAI has 1.1 million GB300s, with a total of 1.44 million GPUs in operation.
⚡

Quick Links

Leaderboards don't tell how models will perform on your data. (Sponsor)

CData ran 22 models against enterprise data through Connect AI. Same correct answer, 178x difference in cost. Read the benchmark.
Build plugins for Claude with the directory submission portal (3 minute read)

Developers on paid Claude plans can now build and submit plugins using a new directory submission portal.
Anthropic Signed an $11.6 Billion Akamai Compute Deal (3 minute read)

Anthropic agreed to spend up to $11.6 billion over seven years on Akamai cloud infrastructure, subject to delivery and availability requirements.
The fastest way to grow on Microsoft Marketplace (Sponsor)

Want to reach the Microsoft Cloud and AI market? Join Frontier Accelerate for Marketplace Access for go-to-market resources, Azure sponsorship, eligibility to participate in co-sell, and certification. Free for eligible partners
Oxford let OpenAI train AI models on Bodleian Library texts (2 minute read)

Oxford allowed OpenAI to use Bodleian Library texts for AI training, sparking concerns about the impact on Oxford's reputation and AI's energy consumption.

Want to advertise in TLDR? 📰

If your company is interested in reaching an audience of AI professionals and decision makers, you may want to advertise with us.

Want to work at TLDR? 💼

Apply here, create your own role or send a friend's resume to [email protected] and get $1k if we hire them! TLDR is one of Inc.'s Best Bootstrapped businesses of 2025.

If you have any comments or feedback, just respond to this email!

Thanks for reading,
Andrew Tan, Ali Aminian, & Jacob Turner


Manage your subscriptions to our other newsletters on tech, startups, and programming. Or if TLDR AI isn't for you, please unsubscribe.