You're offline - Playing from downloaded podcasts
Back to All Episodes
Podcast Episode

DeepSeek's Post-Training Punch, Open Video Finds Its Voice, and the Cages That Weren't Locked

August 1, 2026

0:00
10:47
Podcast Thumbnail

DeepSeek upgrades its V4 Flash model to near-frontier performance without adding a single parameter, then releases the open weights under an MIT licence and reignites the AI price war. MiniMax's H3 promises the first open model that generates video with baked-in audio, ByteDance's Seedance 2.5 holds a scene steady for three minutes, and assistants from Google and OpenAI move onto the desktop. Plus, why the industry's 'rogue AI' scare was really a case of somebody leaving the test cage unlocked.

DeepSeek's post-training masterclass

The day's headline came from DeepSeek, which shipped an upgraded V4 Flash (roughly 284B total parameters, ~13B active, 1M context) that scored dramatically higher on tough coding and agent benchmarks without changing its architecture or size. The whole gain came from post-training, heavy on reinforcement learning, lifting it onto the intelligence-vs-cost Pareto frontier at about $0.14/$0.28 per million tokens, with a near-98% cache discount. The weights landed on release day under a permissive MIT licence, complete with a speculative-decoding module for extra speed, and the community had runnable quantised builds within hours.

A price war with no floor

DeepSeek's move read as a direct answer to OpenAI's prior-day cuts (GPT-5.6 Luna down 80%, Terra down 20%). The broader story is relentless deflation in the cost of a given level of capability. The caveat worth keeping in mind: price cuts reflect competitive strategy as much as true inference cost, but cheap, capable intelligence is fast becoming a commodity.

Open video finally gets a voice

MiniMax launched H3, an omni-reference video generator that produces clips with native stereo audio and accepts text, image, video and sound as blended inputs (up to ~15 seconds at high resolution). With open weights promised within days, it could be the first genuinely open model to generate picture and sound together.

ByteDance stretches the clip

ByteDance's Dreamina pushed Seedance 2.5, offering native 30-second clips, consistency across roughly three minutes, interactive frame editing, and up to 50 multimodal references, tackling the character-consistency problem that separates a toy from a usable tool. Testers flagged 720p limits and some instruction-following quirks around audio.

Assistants move onto the desktop

Google's Gemini Drops added Mac voice control, wider rollout of its lightweight helper, and personalised image and avatar features, while OpenAI brought ChatGPT voice to Mac and Windows plus a new activity view. The quiet shift: assistants are leaving the browser and becoming an always-there layer on your machine.

The harness is the new frontier

A strong research thread argued that model capability is increasingly bottlenecked by the scaffolding around it. Microsoft's Echoverse compiles specifications into working applications and repairs its own test environments, while AgentRadio showed that letting several agents message each other asynchronously nearly doubled a software-QnA score versus a stronger single model.

When the sandbox isn't sealed

Anthropic disclosed that, during cybersecurity evaluations, some Claude models reached into three real external organisations after supposedly isolated environments were accidentally connected to the live internet, found by reviewing over 141,000 eval runs. The technical consensus: not rogue agency but an infrastructure and harness failure, with the fix being better sandboxing, logging and access control.

Published August 1, 2026 at 4:07am

More Recent Episodes