Podcast Episode
Fable 5 Returns, GLM-5.2's Coding Surge, and NVIDIA's 2.42x Faster Text Machine
July 2, 2026
0:00
9:10
Anthropic brings Claude Fable 5 back online with new safety fallbacks, while open coding models like Z.ai's GLM-5.2 close the gap on Western frontier labs. Plus NVIDIA's TwoTower generates text 2.42x faster, Huawei open-sources a 92B model trained on domestic chips, Cognition's security swarm squashes over 1,000 vulnerabilities, and Together AI raises $800M at an $8.3B valuation.
Fable 5 Is Back, With Guardrails
Anthropic has re-enabled Claude Fable 5 after coordinating with the US government, but the relaunch comes with visible safety fallbacks. Updated cybersecurity classifiers may temporarily route flagged requests to Opus 4.8, and biology and chemistry filters remain broad enough to trip on some ordinary queries. The tooling ecosystem responded within hours, with coding platforms and agent orchestrators restoring the model. The more interesting shift is behavioural: builders are increasingly designing multi-model strategies, using a frontier model only for high-value planning while delegating implementation and verification to cheaper models.Open Coding Models Close the Gap
Z.ai isn't just shipping a checkpoint for GLM-5.2, it launched ZCode, a dedicated development environment built around the model. On the benchmark front, GLM-5.2 became the first open model to lead a category on APEX-SWE, posting 55.3% Pass@1 on integration tasks. Analysts caution against declaring open models have overtaken the West, but the coding gap is shrinking fast.NVIDIA's TwoTower Speeds Up Generation
NVIDIA introduced Nemotron-Labs-TwoTower, adapting a 30B model into a diffusion-style language model that writes tokens in parallel using a frozen context model plus a trained writer. The claimed result is 2.42x faster generation while retaining 98.7% of the original quality, all without full retraining.Huawei Goes Open Source
Huawei open-sourced OpenPangu-2.0-Flash, a 512K-context mixture-of-experts model with 92B total and 6B active parameters, releasing weights, inference code and training ops. Notably, it appears to have been trained on Huawei's own accelerators rather than NVIDIA GPUs, a strategically significant proof point under export controls. A larger 505B flagship is planned.Security Swarms and Big Funding
Cognition's Devin Security Swarm uses an Agentic MapReduce pattern to fan bounded agents across a codebase, validate exploitability, and surface confirmed vulnerabilities. A Fortune 500 pilot reportedly found and fixed over 1,000 vulnerabilities in production. Meanwhile, Together AI announced an $800M Series C at an $8.3B valuation, underscoring investor appetite for inference infrastructure.Local Inference Keeps Winning
Local runtimes had a strong day too. A native C++ and ggml implementation generated a 90-minute podcast in under 23 minutes, over four times real-time, and NVIDIA released an NVFP4 quantised Qwen3.6-27B that fits comfortably on 32GB consumer cards. WebGPU Gemma 4 also hit 255 tokens per second on an M4.Published July 2, 2026 at 10:37am