Why I'm Building Muse (2 minute read)
Muse is built as a personal agent that turns vague ambitions into concrete action by planning, emailing, calling, finding resources, and removing friction. Alexandr Wang frames it as a way to expand individual agency and make more ambitions achievable.
|
OpenAI prepares to expand Ultrafast API to more users (1 minute read)
OpenAI appears to be preparing a wider rollout of its Ultrafast API mode. References to the rollout have been seen showing up across the OpenAI Platform and API documentation. The feature was officially previewed with GPT-5.6 Sol. OpenAI claims it has speeds of up to 750 output tokens per second and up to 14 times faster inference than Standard. The mode is powered by Cerebras. Access remains limited to select customers.
|
|
Let's talk about trading compute (16 minute read)
There is an emerging market of compute derivatives. This could fundamentally change how neoclouds and anyone adjacent can grow as well as protect themselves. The problem is pretty urgent for inference clouds. Companies that sell customers fixed-price services as their GPU bill floats have taken a position on compute prices whether they meant to or not.
|
Can AI self-improvement overcome diminishing returns? (38 minute read)
AI is already helping build better AI. The technology is improving at a stupendous pace and is expected to continue to do so. Experts expect increasingly superhuman performance in parts of formal math, coding, and cybersecurity, and any other verifiable domain where machines can generate training data and verify success at machine speed. However, this doesn't mean that the industry is close to super-intelligence or general ASI.
|
Agent (Muse) Compute Demand (5 minute read)
It costs Meta an estimated around 1 to 2 gigawatts of average total power to serve 100 million daily active users. Only around 0.1 gigawatts comes from the GPU/VM layer. 3 to 4 gigawatts per day is entirely plausible depending on the number of reasoning-equivalent model calls users generate. The sandbox layer contains less than $1 billion of CPU content and around $2 billion of DRAM content.
|
|
22 Models, Same Correct Answer, 178x Cost Gap (Sponsor)
Leaderboards don't rank models on how they perform on your data. CData connected 22 models to live CRM, warehouse, and ITSM data through CData Connect AI. When Connect AI carried business context and guardrails, 17 returned the identical correct answer, one model at $0.0009 and another model at $0.157. Read the benchmark
|
Claude computes a nine-loop amplitude in N=4 super-Yang-Mills (18 minute read)
Physicists at Anthropic harnessed Claude, an LLM, to compute a complex nine-loop amplitude in N=4 super-Yang-Mills, a problem once considered computationally infeasible with limited resources. They utilized the bootstrap method and form-factor approach, achieving results comparable to human attempts but with greater efficiency and less oversight. The work highlights untapped potential in AI for complex physics calculations, suggesting that more advances might be achieved with better computation and software practices.
|
|
On Ezra Klein's Podcast With Jensen Huang (40 minute read)
Jensen Huang doesn't believe in superintelligence or that AI will ever be a different kind of thing from software. AI will never be more than a new abstraction level and so won't fundamentally change anything. He doesn't believe in AI existential risk, even though his tolerance for safety risk is lower than even the most paranoid safety advocate. Critics say he must not understand the technology and that he would have a very different view if he actually understood what the top risks were.
|
Elon Musk's SpaceXAI to add another 660,000 AI GPUs this year (2 minute read)
220,000 Nvidia GB300 GPUs will be operational at X's Colossus supercomputer in November. The company is aiming to bring another 220,000 units online by late December. Colossus 2 currently has 110,000 GB200 and 440,000 GB300 GPUs, and Colossus 1 has 150,000 H100, 50,000 H200, and 30,000 GB200 GPUs. This means SpaceXAI has 1.1 million GB300s, with a total of 1.44 million GPUs in operation.
|
|
|
|
|