Claude Cowork and chat are now one Claude (2 minute read)
Claude Cowork and chat are merging into one Claude. Claude Docs and Slides can now produce documents and presentations that users can edit directly, present straight from Claude, or download as a PowerPoint or PDF. The changes will roll out on Pro and Max plans over the next few weeks, with Team and Free plans following soon. Enterprise admins will hear from Anthropic at least 30 days before anything changes for their organizations.
|
OpenAI Expanded ChatGPT Ads with AI Agents (4 minute read)
OpenAI announced Sponsored Agents that let users start conversations with business-sponsored agents after clicking an ad in ChatGPT. It also introduced AI-assisted ad creation in ChatGPT Work, new Ads Manager creative tools, and integrations with HubSpot and Shopify.
|
Your AI agents can now control your Google Home devices (3 minute read)
Google introduced early access to the Model Context Protocol (MCP) for Google Home, enabling AI agents like ChatGPT to control smart home devices. Users must set up a Google Cloud project and provide MCP details to their chosen AI agent for configuration. This update supports all devices in the Google Home ecosystem and is available initially to US subscribers of Google Home Premium Advanced.
|
|
AI Cheating is on the Rise (4 minute read)
Studies have shown that models are cheating on evaluations. The same guardrails preventing models from cheating are likely being used during training, and models may be training to complete tasks that evade these specific guardrails. It's unsurprising that labs occasionally release benchmark results that aren't externally trustworthy. This highlights the value of independent evaluators.
|
HarnessTax (2 minute read)
Language models are changing how software is built and computational problems are solved. The impact of harness choice remains unclear, despite millions of people already using coding agents. An evaluation of 21 model-harness pairs spanning seven models and three harnesses found that harness choice has little effect on task success rate, but can significantly affect the cost. A simple harness can be competitive.
|
How Embedded Evaluators Could Monitor Frontier AI (8 minute read)
Transluce outlined how independent evaluators embedded inside AI labs could investigate risks such as multi-agent coordination, targeted persuasion, evaluation awareness, and concealed reasoning. Proposed approaches included monitoring agent swarms, examining training practices, and testing unreleased models under privileged access.
|
|
Agent Substrate brings high-density, scalable, trusted infrastructure to GKE (9 minute read)
Agent Substrate is now available on Google Kubernetes Engine (GKE). Agent Substrate is an open-source, secure-by-default agent execution runtime engineered to run millions of sandboxes with 10x higher density than standard container runtimes. It delivers sub-500ms resume operations at over 500 suspend/resume activations per second with native zero-trust kernel and network isolation. Agent Substrate runs on any Kubernetes infrastructure and is optimized for GKE.
|
Meta's FLAT for Multimodal Understanding and Generation (7 minute read)
Meta AI introduced FLAT, a method that converts images and text into the same flexible-length sequence of continuous tokens for retrieval and generation. Nested dropout arranges information from coarse to fine, letting models trade computation for visual detail by varying the number of tokens used.
|
Ant Group Released a Finance-Focused Model (6 minute read)
Ant Group released Ling-3.0-flash-Fin, an open-weights model developed with financial institutions for tasks such as source checking, valuation spreadsheets, and report writing. Artificial Analysis reported scores of 23 on its Intelligence Index and 24 on its Finance & Accounting Index.
|
|
Why Salesforce may be AI's adult in the room (4 minute read)
Salesforce announced Koa, a domain-specific AI model designed to enhance business task handling while ensuring data privacy. Built on Nvidia's open model, Koa aims to optimize CRM actions with fewer errors and automate routine tasks, acting as an enterprise empowerment tool rather than replacing jobs. Other AI initiatives include AIFORCE for direct Salesforce instance interaction and CLAUDEFORCE, which enhances sales capabilities with pre-built skills.
|
Introducing the DeepMind Institute (3 minute read)
Google DeepMind launched the DeepMind Institute to study AGI's technical and societal implications across safety, governance, institutions, and human values. Led by Demis Hassabis, James Manyika, and Shane Legg, it will convene interdisciplinary researchers from inside and outside Google.
|
|
|
|
|