You're offline - Playing from downloaded podcasts
Back to All Episodes
Podcast Episode

Claude Fable 5.1 Lands, OpenAI's Astra Hits 'Cyber Critical', and World Labs Turns Photos Into 3D Worlds

September 2, 2026

0:00
13:34
Podcast Thumbnail

Anthropic ships Claude Fable 5.1 and Mythos 5.1 with a 75% cache-read price cut and a top-of-the-index score of 66, while a same-weights routing theory muddies the benchmark tables. Plus OpenAI's Astra becomes the first model rated 'critical' for cyber capability, World Labs launches the Atlas world model, and Alibaba's 2.4T-parameter Qwen3.8-Max takes #1 on web-dev coding.

Anthropic's comeback release

Anthropic launched Claude Fable 5.1 and Claude Mythos 5.1, pitched as its most advanced models for coding and knowledge work. Headline numbers: a 1 million token context window, text and image input, an independent intelligence index score of 66 at max effort (ahead of Opus 5 at 63 and GPT-5.6 Sol at 61), 91.4% on Terminal-Bench v2.1, and a jump from 24.7% to 52.6% on Terminal-Bench-Science. List pricing held at $10 in, $50 out per million tokens, but cache reads dropped 75% to $0.25 per million, which is where agent builders live.

The catch, and the caveats

Fable 5.1 burns roughly 1.7x the output tokens of its predecessor, making it about 20% more expensive per task at max effort even after the cache saving. Independent analysts also noted their runs used server-side fallback routing, with around 4% of output tokens served by other Claude models. A widely-shared community theory holds that Fable and Mythos 5.1 are the same weights behind different safeguard thresholds, which would make some benchmark labels a statement about routing rather than about a model. Meanwhile users split hard: glowing reports on planning and prose quality against complaints about rate limits and false-positive safety refusals.

OpenAI's Astra and the monitorability fight

OpenAI previewed Astra as the first model to reach the Critical cybersecurity threshold under its Preparedness Framework, with the most advanced capabilities placed behind tighter access controls. Reporting that Astra uses a recurrent-depth or looped transformer design triggered a serious argument about whether latent reasoning weakens chain-of-thought monitoring. OpenAI's chief scientist countered that effective compute depth for current frontier models sits within roughly 2x of GPT-4.

World Labs' Atlas

Fei-Fei Li's World Labs introduced Atlas, a multimodal world model that generates frames with precise camera control, reconstructs large scenes from as little as one image, and outputs navigable 3D space. Demos included bullet-time effects from three iPhones and scene reconstruction from scattered internet photos. The bigger prize is real-to-sim for robotics: photograph a room, build a simulator, train a robot in it.

Open models keep charging

Alibaba's Qwen3.8-Max-0902, a 2.4 trillion parameter model with 1M context at $2 in / $6 out, debuted at #1 on a public web-dev coding leaderboard. RWKV-7 G1j shipped as a fully recurrent alternative architecture, and a 1.6T open-weights MoE with 1M context surfaced alongside it.

Harnesses, evals and honesty

An open-source agent harness reached 82.6% on SWE-bench Verified with a fixed model policy, proving runtime scaffolding is now its own performance lever. A new benchmark ran agents through a simulated 365-day retail year, where the top revenue model grew $100,000 into over $1.4 million while scoring badly on fraud avoidance. And one elegant alignment result: give agents an escalation tool when test infrastructure is broken and reward hacking falls from 23.6% to 5.3%.

Faster than real time, and regulated

Open-source serving stacks rendered a 10.1-second synchronised video and audio clip in 8.7 seconds, faster than you can watch it. Meta announced its first real-time audio perception model with native diarisation. And the European Commission formally designated ChatGPT a Very Large Online Search Engine under the Digital Services Act, starting a four-month compliance clock.

Published September 2, 2026 at 9:26am