Podcast Episode
Frontier Model Day: Grok 4.6, Qwen 3.8-Max Open Weights, DeepSeek V4 Pro's Price Shock, and the Great Chain-of-Thought Heist
August 13, 2026
0:00
13:11
A blockbuster day in AI: xAI's Grok 4.6 hits frontier performance at $2/$6 per million tokens, Alibaba open-sources a 2.4T-parameter Qwen3.8-Max, and DeepSeek V4 Pro lands at roughly 57x cheaper than rival frontier models. Plus Microsoft's first from-scratch reasoning model MAI-Thinking-1, Lightricks' LTX-2.5 video model with native audio, Google DeepMind's SL2T sign-language translation on Android, and a startling paper showing proprietary 'encrypted' reasoning traces can be extracted straight out of commercial APIs.
Grok 4.6 Redraws the Price-Performance Map
xAI shipped Grok 4.6, and independent evaluation puts it at 61 on the Intelligence Index, roughly level with GPT-5.6 Sol Max, with 88.4% on Terminal-Bench v2.1 and 1753 Elo on GDPval-AA v2. The headline isn't the score, it's the sticker: $2 per million input tokens and $6 per million output, materially below frontier peers. xAI credits a longer supplemental training run, regenerated SFT traces, and agentic reinforcement learning across coding, web, CAD and kernel optimisation. Elon Musk says Grok 4.7 has already finished initial training.Qwen3.8-Max Goes Open Weight at 2.4 Trillion Parameters
Alibaba released Qwen3.8-Max as open weights: a 2.4T total, 95B active mixture-of-experts model with 92 layers, 512 experts, hybrid Gated DeltaNet and Gated Attention blocks, and a native 262K context extendable towards 1M. vLLM shipped day-zero support with 4-bit checkpoints for NVIDIA B300 and AMD MI355X; Together AI and Baseten followed. Caveat: the initial drop is text-only. At bf16 the weights run to roughly 5TB, though Unsloth claims a dynamic 1-bit quantisation squeezing it to 397GB. A 27B variant is promised this week.DeepSeek V4 Pro and the Price War
DeepSeek's V4 Pro reached general availability at around $0.435 per million input and $0.87 per million output — described by Cline as roughly 57x cheaper than Fable 5, with a 15.8% Terminal Bench gain over the preview. Capability reactions were mixed, suggesting future gains may hinge more on RL environments than raw scale.Microsoft's First From-Scratch Reasoning Model
Mustafa Suleyman announced MAI-Thinking-1, Microsoft's first reasoning model built from scratch, now in Foundry. The team's first request was feedback on tool use — a tell that this is aimed at applied agentic work rather than leaderboards. Upstage's Solar Pro 4 also leapt from 14 to 42 on the Intelligence Index.Open Video, Vision and Voice
Lightricks' LTX-2.5 landed in Diffusers with joint video and 48 kHz audio generation, prompt-controlled clip length, a two-pass quality mode, tile rendering for lower memory, plus native multishot generation and Diffusion Fidelity Rendering that allocates compute by scene complexity. Cohere launched North Micro Vision, an Apache-2.0 small document-understanding VLM. Deepgram's Flux TTS claims roughly 80 ms response time with mid-call adaptation.Google DeepMind's SL2T Brings Sign Language to Android
SL2T translates American Sign Language to text as a native input method on Android and Pixel 11, with body pose tracking on-device, translation server-side, and explicit optimisation for one-handed signing.Stealing Reasoning Traces
A paper claims proprietary encrypted chain-of-thought blobs from major APIs can be replayed and reconstructed by weaker sibling models, with recovered token counts matching billed thinking tokens. One decoded trace shows a model recognising a competition maths problem from memory rather than deriving it. The vulnerability is reported patched.Efficiency Wins
Snowflake's 4B SQL autocomplete model beat their previous 30B mixture-of-experts, cutting median latency 71%. Expedia's move to Keras 3 delivered 30% faster training and 70% lower inference latency. And Google's ResidencyRL work lifted diagnostic accuracy under adversarial conditions from 81% to 88% across nearly 50,000 simulated telehealth encounters.Published August 13, 2026 at 2:42am