Introducing H3 Max by fal (5 minute read)
H3 Max is a post-trained version of MiniMax H3 optimized for maximum speed. It can generate a 5-second-long video in under 3 seconds. The model is now available on fal for 50% off for the first week. This article discusses how the model was trained.
|
Introducing Parse: Enterprise document intelligence at scale (5 minute read)
Cohere Parse converts complex, multimodal files into structured machine-readable data. A cost-effective vision language model for processing large volumes of enterprise documents, it detects and understands key visual elements and can process documents and images across nine major world languages. Parse can be accessed through the Cohere API for just $1.50 per 1,000 pages. A free version is available in Cohere Space for users to try.
|
|
GPT 5.6 Discounts & Jevons Paradox (4 minute read)
OpenAI introduced large discounts on some of its models between July 27 and August 14. This resulted in Luna token usage jumping 13.8x and Terra token usage rising 5.6x while Sol, which remained at list price, only saw a bump of 1.1x. Most of the share gained by OpenAI's discounts came from other labs as opposed to cannibalization within the OpenAI family of models. Nearly a third of users who tried a discounted OpenAI model kept using it after the discounts expired.
|
Accelerating MiniMax-H3 (12 minute read)
A detailed benchmark of MiniMax-H3 video generation on 8× H200 GPUs showed SGLang Diffusion reaching 1.95x lossless speedups and as much as 6.24x with step reuse and sparse attention.
|
AI and the city (8 minute read)
AI is driving more distributed business formation in the US, with a rise in new companies forming outside major cities and in outer suburbs. This trend began during COVID-19 with remote work and accelerated with AI, yet business formation does not correlate directly with work-from-home levels. AI labs and product companies still cluster in major metro areas like San Francisco, highlighting the continued value of agglomeration for innovation.
|
|
Support persistent reasoning effort (2 minute read)
Codex has added 'persistent' to the reasoning-effort protocol and TypeScript SDK types. When a user targets a custom Responses-compatible provider whose model-defined effort is literally persistent, this now deserializes to the new Persistent variant and is unconditionally rewritten to disabled. Previously, it remained Custom("persistent") and was forwarded unchanged. Existing configurations and resumed sessions can therefore send a different value and change behavior or be rejected, so this alias should be limited to providers that define the translation.
|
Gemini Omni 1.1 Flash (4 minute read)
Google introduced Gemini Omni 1.1 Flash with new controls for extending scenes, interpolating first and last frames, 4K upscaling, and faster video iteration through the Gemini API.
|
|
Nvidia Climbing the Wall of Worries (7 minute read)
The average sell-side estimates for Nvidia's FY 2028 revenue a year ago were around $310 billion. Nvidia guided for around 70% revenue growth in FY 2028 yesterday, which translates to approaching around $700 billion in revenue next year. While analysts kept updating the estimates throughout last year, they were still short by around $125 billion. Nvidia's revenue growth of around 70% could be the floor for next year.
|
An update on AI's most important number (8 minute read)
OpenAI and Anthropic are seeing unprecedented growth, with revenues surpassing $100 billion combined, fueled by rapid adoption and technical advancements. This swift rise challenges expectations, as typical tech growth slows at such scales, potentially reshaping the economy if sustained.
|
|
| | |