You're offline - Playing from downloaded podcasts
Back to All Episodes
Podcast Episode

Anthropic's Agent Platform Goes GA, AT&T Routes 40% of AI to Open Models, and OpenAI Powers Up NVIDIA's Vera Rubin

August 21, 2026

0:00
13:14
Podcast Thumbnail

Anthropic moves computer use, the browser tool, Skills and Files to general availability while OpenAI turns the ChatGPT desktop app into an action-taking assistant with an Apple Messages plugin. AT&T reveals it already routes 40% of employee AI usage to open models, cutting coding costs 56% for a 2% quality drop. Plus Meta's Muse Spark 1.2 benchmark sweep, OpenAI's first Vera Rubin racks, one-shot robot learning from Generalist AI, and a 1B-parameter model trained from scratch for $250.

Anthropic's Agent Stack Grows Up

Anthropic has moved a large slice of its agent platform into general availability: computer use, the browser tool, the Skills API and the Files API are all now production-ready on the Claude Platform. Skills add versioned, reusable procedures, while Files gains expiration control, a 5x rate-limit increase to 500 RPM and 1 TB of storage per organisation. An AG-UI adapter for Claude Managed Agents lets developers stream text, tool calls and thinking straight into their own interfaces.

OpenAI Turns the Desktop Into a Workspace

OpenAI shipped an Apple Messages plugin for ChatGPT Work and Codex on Mac, letting the assistant search, catch up on, draft and send messages. ChatGPT Sites gained collaborative editing with Codex handling git and CI, alongside shared read-only conversation links and PR-context sharing. Transparent backgrounds arrived in preview for GPT-Image-2, and Computer History plus cross-app memory rolled out to the EEA, UK and Switzerland for Pro, Business and Enterprise Mac users.

The AT&T Datapoint

The most strategically loaded number of the week came from AT&T's internal deployment: 40% of employee AI usage already routes to open models, with a target of 60–70%. Coding costs are down 56% for a 2% quality drop, at 45 billion tokens a day. Meanwhile Gemma passed 1 billion downloads and Kimi K3 rolled out to over half of Ollama's subscriber base with US and EU hosting and zero data retention.

Muse Spark 1.2 Sweeps the Arenas

Meta's Muse Spark 1.2 posted a strong multimodal and agentic showing: a 2.1% net improvement in Agent Arena, up from 0.9% in v1.1, with Bash Recovery up 11.4%. Design-focused evaluations placed it first for Video-to-Website and second for Image-to-HTML, sitting on the price-preference Pareto frontier. Meta also previewed WildArtifactBench, an internal evaluation for practical multimodal tasks.

Rubin Racks and Frontier Pretraining

OpenAI's first NVIDIA Vera Rubin racks are installed and running the training stack, explicitly tied to next-generation frontier pretraining. It is a rare, concrete signal about the physical scale-up behind the next round of models.

One-Shot Robots

Generalist AI introduced GEN-1.5, a one-shot learner for embodied AI. Show the robot a task once and it reproduces and generalises the behaviour. Observers called it the GPT-2 moment for robotics. Elsewhere, DaxAI debuted an all-terrain robot-horse claiming 300 kg payload and 40 km/h.

Harnesses, Memory and Honest Negative Results

Chroma launched Foundation, a research preview of agent memory built from prior sessions. A related paper on harness continual learning shows prompts, skills and routing rules can evolve independently of model weights, with guarded harness evolution reporting over 10% gains. A counterweight study found memory-based self-improving agents look far less impressive once task order and evaluation variance are controlled for.

Small, Cheap and Specialised

A developer trained a 1B-parameter Kimi-K3-style mixture-of-experts model from scratch on a single H200 for roughly $250, beating GPT-2's 124M on HellaSwag. And users report Qwen3.8-27B has lost factual recall versus 3.6 while gaining coding and tool-use skill, a deliberate trade of memorised trivia for agentic competence.

Published August 21, 2026 at 7:24am