Podcast Episode
The Stealth Model Nobody Will Claim: Ox Alpha, DeepSeek Goes Multimodal, and Nvidia's $6B Poolside Deal
August 22, 2026
0:00
11:45
A mystery coding model called Ox Alpha appears out of nowhere and posts 80% on software engineering tasks, while DeepSeek ships V4-Flash-Vision-Exp with a new Files API. Plus OpenAI cuts GPT-5.6 Sol pricing by over 20% as Codex hits 20 million users, Nvidia pays Poolside roughly $6B for its model factory, and NVIDIA AVO claims 100% on the public ARC-AGI-3 environments.
It was billed as a quiet day. It wasn't.
The stealth model nobody will claim
A model called Ox Alpha turned up in coding tools with no announcement and immediately started posting numbers: roughly 80% on a set of software engineering tasks, against about 65% for Fable and 52% for GPT-5.6 Sol. One developer reported it finding two real bugs in code already audited by other frontier models. Speculation converged on the GLM family from Zhipu, likely an efficient flash or vision variant rather than a giant new base model, based on its speed profile and its weaker prefill behaviour. It feeds a wider thesis: post-training and infrastructure now matter more than raw parameter count.DeepSeek adds eyes
DeepSeek shipped V4-Flash-Vision-Exp, adding multimodal input while reportedly preserving the text performance of V4-Flash. Images are billed at up to 384 tokens each at Flash pricing, and a new Files API lets you upload an image once and reference it by ID. DeepSeek claims multimodal agent performance close to Opus-4.8, with 83.9 on Terminal Bench 2.1 and 75.9 on Toolathlon-Verified.Prices down, usage up
OpenAI cut GPT-5.6 Sol API pricing by more than 20% for three months, stacking with product promotions elsewhere. Codex reportedly reached 20 million active users, and OpenAI added per-API-key spend tracking with hard monthly limits, a sign that agentic workloads have become genuinely unpredictable.Nvidia buys the factory, not the company
Nvidia will reportedly pay Poolside around $6B to license its "Model Factory" development stack, invest a further $1B at a $12B pre-money valuation, and extend offers to 109 employees behind the Laguna coding model. Reaction split between optimism for US open models and worry that Laguna's independent roadmap ends here.Benchmarks, caveats and harder tests
NVIDIA AVO cleared all 183 levels across 25 public ARC-AGI-3 environments with no instructions or stated goals. François Chollet noted this is the public demonstration set, not the private benchmark. Meanwhile new evaluations are getting brutal: SWE-bench Science puts top agents under 50% pass rate, and CADBench sits at 24.6%.Big models, small machines
UC Berkeley's FreeToken reportedly runs a 753B-parameter GLM model at 14.9 tokens per second on a single RTX Pro 6000, and a 35B model at 39.3 tokens per second on an 8GB laptop GPU, claiming two to four times the throughput of common local runtimes. Separately, Percy Liang announced the fully open Marin 535B-A23B training run has begun, targeting 18.75 trillion tokens.Recycling thoughts, and robots that feel
A DeepMind paper on Recirculation feeds deeper-layer activations back into earlier processing at inference time with no retraining, reporting 60% fewer contextualisation errors and a 21% maths gain. Pandora's Router reframes model routing as a search problem with costly inspection. And Jim Fan introduced T-Rex, a tactile-reactive manipulation stack with what's described as the largest open tactile dataset yet: 50 hours and around 5,500 episodes on 22-degree-of-freedom hardware.Published August 22, 2026 at 8:49am