You're offline - Playing from downloaded podcasts
Back to All Episodes
Podcast Episode

Grok 4.5 Undercuts the Frontier, GPT-Live Talks Back, and the Great AI Price War

July 9, 2026

0:00
10:23
Podcast Thumbnail

xAI ships Grok 4.5, a 1.5-trillion-parameter coding-and-agents model trained with Cursor at roughly a quarter of Opus pricing. OpenAI launches GPT-Live, a full-duplex voice model, and teases GPT-5.6 Sol. Plus Cognition's blazing SWE-1.7, Mistral's one-camera robot brain, NVIDIA and LangChain's open agent stack, ByteDance's Seedream 5.0 Pro, and a new voice-cloning champion.

Grok 4.5 makes cost the headline

xAI has publicly launched Grok 4.5, its first model trained specifically for coding and agents, built in partnership with Cursor. At 1.5 trillion parameters it is roughly 3x larger than Grok 4.3, yet the story isn't raw size, it's economics. Pricing lands at $2 per million input tokens and $6 per million output tokens, with cache hits discounted to $0.5. Independent evaluations from Artificial Analysis place it at #4 on their broad intelligence index and #4 on agentic evals, while burning dramatically fewer tokens than Opus 4.8 or GPT-5.5 on the same coding tasks. The takeaway from analysts and developers alike: not the best model overall, but arguably the best buy, and a direct shot at competitor pricing.

OpenAI ships GPT-Live and teases Sol

OpenAI pre-announced GPT-5.6 Sol, alongside Terra and Luna, for a Thursday launch, with early testers unusually bullish on maths, coding, and computer use. But the thing that actually shipped is GPT-Live, a third-generation full-duplex voice architecture. Unlike turn-based voice pipelines, it listens and speaks simultaneously, and quietly delegates web search and deeper reasoning to a frontier model in the background. It's live in ChatGPT across web, iOS, and Android.

Cognition, Mistral, NVIDIA and the efficiency wave

Cognition launched SWE-1.7, built on a Kimi K2.7 base, claiming near-frontier coding at around 1000 tokens per second and just $1.97 per task, with self-compaction for long jobs. Mistral released Robostral Navigate, an 8B embodied navigation model that steers a robot using a single ordinary RGB camera and claims state-of-the-art results on the R2R-CE benchmark. NVIDIA and LangChain unveiled the NemoClaw Deep Agents Blueprint, a fully open enterprise agent stack claiming benchmark-leading results at roughly 10x lower inference cost, an aggregate score of 0.86 at $4.48 versus $43.48 for the closest rival.

Media models and voices

ByteDance rolled out Seedream 5.0 Pro, an image model built for design-sensitive work: text rendering, layout, separate layers, multilingual text, and structured outputs. And Artificial Analysis launched a Controlled Voice Arena for standardised voice cloning across 8 shared voices, crowning Cartesia Sonic 3.5 the overall leader at 1122 Elo, with Fish Audio S2 Pro leading the open-weight pack at 1034 Elo.

The metric that matters

Threaded through nearly every story is one industry-wide shift, echoed by engineers at Databricks and beyond: the number that matters now is dollars per task, not dollars per token. When a fast, cheap model can finish the job while spending a fraction of the tokens, single-provider loyalty weakens and routing between models becomes the real unlock. There's also chatter that China's MiniMax plans to open-source a roughly 2.7-trillion-parameter model, another sign that frontier capability is diffusing fast across price tiers.

Published July 9, 2026 at 7:47am

More Recent Episodes