You're offline - Playing from downloaded podcasts
Back to All Episodes
Podcast Episode

OpenAI's 8GW Ohio Bet, Stripe's $7B OpenRouter Deal, and Cursor Ships Its Own GitHub

August 18, 2026

0:00
11:32
Podcast Thumbnail

OpenAI goes vertical on power with a 4+ GW NVIDIA commitment and an 8 GW Ohio campus, while Stripe reportedly buys model-routing layer OpenRouter for over $7B and Cursor launches Origin, its own code hosting platform, mid-GitHub-outage. Plus Cartesia's Sonic 3.6 takes #1 on both voice leaderboards at 136.1 chars/sec, NVIDIA's Nemotron 3.5 Lightning packs 30B parameters with 3B active, and new research suggests memory and compaction, not size, are this week's real capability multipliers.

OpenAI buys the power, not just the chips

OpenAI's compute strategy has gone very literal. Mark Chen pointed to a commitment of more than 4 GW of NVIDIA capacity, alongside an 8 GW campus in Ohio built and operated by SB Energy, with NVIDIA backing the initial 4.25 GW and a buildout running through 2032. First 800 MW is expected around 2028. The notable part isn't the scale, it's the vertical coupling: power, land, data centres, chips and long-dated access all locked together.

Stripe's $7B bet on the routing layer

Stripe has reportedly agreed to acquire OpenRouter for over $7B, a striking outcome for an aggregation layer that takes roughly 5% of spend. The obvious question is margin durability, and it's already being answered: OpenRouter cut GPT-5.6 Sol pricing while Vercel did the same on its AI Gateway. Model brokerage is becoming a pricing battlefield rather than a stable tollbooth.

Cursor launches Origin

Cursor shipped Origin, a repository hosting product built directly into the editor covering repo management, PRs, review and deploy integrations, with GitHub sync retained. It landed during a major GitHub outage. The strategic read: agentic coding tools want the full loop, not just autocomplete.

Cartesia takes both voice leaderboards

Sonic 3.6 hit #1 on both the Provider Voice and Controlled Voice leaderboards, with claimed naturalness gains across 44 languages and throughput measured at 136.1 characters per second, materially faster than several premium rivals. Latency is what separates a conversation from an interrogation.

Nemotron 3.5 Lightning: efficiency at the architecture level

NVIDIA's Nemotron 3.5 Lightning is a 30B mixture-of-experts model with only 3B active, trained for high-throughput agent execution, with multi-token prediction feeding speculative decoding plus quantised checkpoints and drafters. Inference efficiency is moving beyond quantisation into architecture.

Memory and compaction as the new scaling track

A 150M-parameter model called BDH-CQ, doing latent-space reasoning with temporary memory, reportedly hit 29.5% pass@2 on ARC-AGI-1 at around $0.0007 per task. Separately, OpenAI reported that retained reasoning plus compaction lifted GPT-5.6 Sol from 13.3% to 38.3% on ARC-AGI-3 using roughly 6x fewer output tokens. Better and cheaper together is rare.

A provocative claim about reinforcement learning

A paper introducing ReasonMaxxer claims RL for reasoning alters only 1-3% of token positions, clustered at high-entropy decision points, with the promoted token already in the base model's top 5. It replicates the gains without RL at roughly 1000x less compute. Readers pushed back hard on the "always top 5" claim, but even the soft version reframes RL as sparse reranking rather than new capability.

Skills are recipes, not encyclopaedias

Related work found agent skills help mainly through procedural anchoring (65.7%) rather than factual injection (4.5%), with precision collapsing as skill pools grow. And in the accidental-discovery category, MiniMax H3, released as a video model, is turning out to be an unusually strong still-image generator.

Published August 18, 2026 at 12:22am

More Recent Episodes