You're offline - Playing from downloaded podcasts
Back to All Episodes
Podcast Episode

Qwen Opens Its Biggest Model, a 10-Million-Token Memory, and a Microscope for AI

August 5, 2026

0:00
10:49
Podcast Thumbnail

A deceptively 'quiet' news day delivered a flood of releases. Alibaba's Qwen 3.8-Max goes open at the top tier alongside a 27B sibling that fits on a single consumer GPU, NVIDIA open-sources a reasoning model for self-driving cars, and Mistral ships a 3B on-device safety guard. Rounding it out: a 10-million-token context model, ternary-weight efficiency running at 200+ tokens/sec on a Mac Mini, Cursor's faster training kernel, and Goodfire's new interpretability platform Silico.

Qwen Opens Its Biggest Model — and a Tiny One You Can Run at Home

Alibaba's Qwen team unveiled Qwen 3.8-Max, a roughly 2.4-trillion-parameter model billed as "better and cheaper" and, for the first time on a Max-tier model, promised to open the weights. The launch leaned on agentic feats rather than chatbot benchmarks, including claims of 10-day autonomous software builds and a closed-loop chip-design pipeline that shrank a circuit from over 8,000 logic gates to under 700. Alongside it came Qwen 3.8-27B, a compact sibling that Unsloth's Daniel Han says should fit in about 17GB of VRAM, putting frontier-adjacent quality within reach of a single consumer graphics card.

NVIDIA Opens a Brain for Self-Driving Cars

Jensen Huang announced Alpamayo 2 Super, an open reasoning model built specifically for autonomous vehicles and released for commercial use. NVIDIA's pitch is provocative: open models are a safety enabler, because a shared, inspectable driving system can be scrutinised by far more engineers than any single carmaker's secret stack.

Mistral's Pocket-Sized Safety Guard

Mistral launched Shieldstral, a 3B open-weights safety and moderation model designed to run on-device. It scores content in a single forward pass, handles 12 languages and accepts images as well as text, letting developers embed moderation directly into apps instead of routing everything through the cloud.

A 10-Million-Token Memory

Pokee AI released Pokee-Isaac, a 28B model claiming a 10-million-token context window with strong accuracy at length, deployable from a single high-end GPU. If it holds up, it points to models that can hold entire codebases or years of records in working memory at once.

Ternary Weights and 200 Tokens a Second on a Mac Mini

deepgrove's Maple-Preview is a 20B reasoning model built on ternary weights, compressing each internal value to essentially -1, 0 or +1, and reportedly hitting 200+ tokens per second on a Mac Mini. It's a striking example of the field chasing efficiency, not just scale.

Cursor Open-Sources "Mixture-of-Kittens"

Cursor released MoK, a training megakernel that fuses mixture-of-experts communication and compute into a single operation, claiming up to 2.37x speedups over strong public baselines. Faster training means cheaper, and greener, models downstream.

Goodfire's Silico: A Microscope for AI

Goodfire launched Silico, a platform for interpretability, the science of looking inside models to understand what they're actually doing. Researchers immediately applied it across language models, medical imaging and biology. As we lean on AI for high-stakes decisions, tools that let us inspect the "black box" may prove among the most important work in the field.

Published August 5, 2026 at 1:38am

More Recent Episodes