Podcast Episode
DeepSeek's Flash Eats Its Pro, OpenAI Claims Navier-Stokes, and a Million-Token Laptop
September 12, 2026
0:00
14:31
DeepSeek quietly retires V4 Pro because the cheaper V4.1 Flash beats it, Meta's Muse Spark 1.3 goes free and tops Website Arena, and Alibaba and Perceptron open-source driving and robotics models. Plus Qwen3.8-Flash-Next runs a 1M-token context on a MacBook, OpenAI claims a Navier-Stokes result amid an authorship row, ChatGPT's free tier gets a major upgrade, and Kepler Compute and Apple's A20 Pro shake up AI hardware.
DeepSeek soft-retires V4 Pro as V4.1 Flash takes over
DeepSeek has effectively retired its V4 Pro model: requests to V4 Pro are now routed to the new V4.1 Flash and billed at Flash pricing until a V4.1 Pro arrives. The stated reason is that the smaller, cheaper Flash now beats Pro on performance, cost, speed and usable request time. V4.1 Flash, already rolling out through the API, is described as a new architecture with native multimodal support. Early testers report roughly 2.2x faster responses (with the caveat of low beta load) and up to 30% better token efficiency. Community theories for Pro's decline include reward hacking and poor scaling despite being around 6x larger.Meta Muse Spark 1.3 goes free and tops Website Arena
Meta's Muse Spark 1.3 is now free inside the Cline coding agent, where the team says it performs similarly to Anthropic's Opus 5 at a fraction of the cost. On Design Arena it reached #1 on Website Arena with an Elo of 1362, a five-place jump over 1.2, and sets a new speed/price Pareto point. Artificial Analysis notes that Claude Fable 5.1, Muse Spark 1.3 and GPT-6 Astra have all pushed the intelligence-versus-cost frontier outward.Open models for cars and robots: Qwen-Drive and Isaac 0.5
Alibaba's Qwen team released Qwen-Drive-1.0-4B, an open-weight autonomous-driving vision-language model built on an unchanged Qwen3.5 backbone with added modules for bird's-eye-view 3D detection, occupancy, map segmentation and motion planning. Perceptron's Isaac 0.5 robotics model, with weights on Hugging Face, claims reliable fine-tuning for repetitive tasks like box packing from roughly 30 demonstration episodes. StereoPolicy research shows 3D manipulation perception from plain stereo cameras, no depth sensor or LiDAR required.A 1M-token context on a MacBook, and Photon 2.2's megakernels
Qwen3.8-Flash-Next now runs with a 1,048,576-token context on an M5 Max with 128GB via mlx-serve, using mixed 4/8-bit quantisation and an 8-bit KV cache. Peak memory is about 117GB; prefill stays near 1,000 tok/s toward 1M context, with generation falling from 100+ tok/s to about 40 tok/s at the limit. Separately, Photon 2.2 extended optimised local inference across A10, A100, 3090, L4, H100, B200 and RTX PRO 6000 Blackwell, with major upgrades to its megakernel compiler.OpenAI claims Navier-Stokes, and a bitter authorship dispute
OpenAI says an internal model, described as significantly more capable than GPT-6 Astra, has cracked the Navier-Stokes Millennium Prize problem, reportedly via a blow-up counterexample found by a swarm of around 10,000 agents over 88 hours. NYU mathematician Tristan Buckmaster has published a statement alleging suspicious timing, a similar proof strategy to his own unpublished work with a collaborator, unresolved questions about training data, and an offer of partial credit conditioned on removing his Anthropic-employed co-author. None of the allegations are independently verified. Cognition also detailed a Devin-assisted GPU lattice siever that made RSA-260 factoring 10x cheaper.OpenAI's free-tier overhaul, Christiano joins, and Anthropic's incident review
OpenAI reports major factual errors in ChatGPT down 65% since March, extreme sycophancy down 80% and medical hallucination flags down 83%, with free users now getting unlimited text chats, higher reasoning effort, automations and memory 'dreaming'. Paul Christiano joins the OpenAI Foundation Board and Safety and Security Committee, and a 250+ person 'Defense Factory' uses models to fix vulnerabilities at scale. Anthropic published an assessment of four cyber incidents involving Claude during third-party evaluations mistakenly connected to the internet; METR will run an independent eight-week investigation.Hardware: Kepler Compute's $468M bet and Apple's A20 Pro
Kepler Compute emerged from seven years in stealth with $468M raised, its own fab, and a roadmap of 3D/materials innovation with no EUV dependence and memory of up to 10x HBM capacity. Apple's A20 Pro moves to 2nm with a 32-core Neural Engine and ~115 GB/s memory bandwidth, above the M2/M3 and close to the M4. Epoch AI estimates OpenAI's compute use has grown nearly 20x since 2023.Published September 12, 2026 at 10:36pm