You're offline - Playing from downloaded podcasts
Back to All Episodes
Podcast Episode

Qwen3.8-27B Goes Local-Frontier, OpenAI Pauses Frontier Training, and Mojo Goes Open Source

August 19, 2026

0:00
13:51
Podcast Thumbnail

OpenAI hits pause on its biggest frontier training run to harden safety and monitoring, while a 27B open model from Alibaba starts trading blows with systems many times its size on consumer GPUs. Plus Z.ai's GLM-5.3, Modular open-sourcing Mojo, NVIDIA's two-command TensorRT deployment, Cerebras CS-4, and a DRAM price spike that's making local AI hardware painful.

OpenAI Pauses Its Biggest Frontier Run

The day's most significant story wasn't a launch, it was a stop. OpenAI paused some frontier reinforcement learning training for two weeks and is still holding its largest planned frontier RL run while it strengthens workload and network isolation, continuous security testing, and multistage monitoring. Sam Altman framed it as capabilities outpacing safety readiness; Greg Brockman said confidence in safety will increasingly set the pace of scaling. Notably, monitoring reportedly adds roughly 20% overhead, and sampled-token monitors can page safety teams within about 30 minutes. Models already near shipping are unaffected.

Qwen3.8-27B Becomes the Local Frontier Moment

Alibaba's Qwen3.8-27B dominated the open-model conversation, hitting #1 local model in Cline within four days, #7 on Artificial Analysis' Agentic Index, #6 among open-weight models on Vals Index v2, and #1 on Harvey's legal benchmark among open weights. Independent benchmarking showed Q4_K_M quantisation retaining roughly 99.97% of Q8 quality while saving around 10GB, and users are running it inside 16GB of VRAM at 73k context. Sceptics push back that it still trails Opus 5 on long, messy real coding work, and that arcade-game clone demos reward memorisation.

GLM-5.3: A Post-Training Story

Z.ai shipped GLM-5.3 via API at the same price as GLM-5.2, tying Kimi K3 at 60 on the Intelligence Index with a 246-point jump on GDPval-AA v2 to 1770 Elo, all on the same 753B total / 40B active MoE footprint and 1M context. The gains appear to come from post-training: asynchronous RL, executable sandbox training, and on-policy distillation to avoid catastrophic forgetting.

Mojo Goes Open Source

Modular open-sourced the Mojo language under Apache 2.0 and positioned its platform as a portability layer across accelerators, including Qualcomm datacentre AI chips. Toolchain openness plus hardware abstraction arriving together is the real story.

Faster Decoding, Everywhere

NVIDIA launched TensorRT Model Connect in public preview, converting supported Hugging Face models to end-to-end TensorRT inference in two commands with no intermediate ONNX export. DFlash 2 claims Qwen3.8-27B at 70 tok/s on an M5 Max with up to 4.6x faster autoregressive decoding. Cerebras announced CS-4 with claims of 10 trillion parameter models at 1000 tok/s. And Alibaba's XuanTie C950, a 64-core RISC-V server CPU, reportedly runs Qwen3.8-27B at about 30 tok/s with no GPU at all.

Memory Prices Bite

DRAM has climbed roughly 500% in twelve months, with 128GB DDR5 kits listed at $3,399 and 64GB ECC RDIMMs moving from about $300 to $1,800. Workstation GPUs are following, with the RTX Pro 6000 MSRP appearing at $19,999.

Measuring AI in Public

Researchers at MIT, Stanford and elsewhere launched the Public AI Observatory: 24,521 consented conversations, 52 models, nearly 100,000 turns and 145 labelled features spanning 2023 to 2026, deliberately independent of vendor self-reporting. Separately, a study of 1,902 multi-agent coding runs found that naming a coordinator doesn't reliably help, and swapping repeated one-to-one messages for shared files cut output tokens by about 42% at eight agents.

Published August 19, 2026 at 10:33am

More Recent Episodes