You're offline - Playing from downloaded podcasts
Back to All Episodes
Podcast Episode

Open-Weight Deluge: Tencent Hy3, Meituan's 1.6T LongCat, and Anthropic Finds Claude's 'Working Memory'

July 7, 2026

0:00
9:45
Podcast Thumbnail

A supposedly quiet week delivered three enormous open-weight models in days: Tencent's 295B Hy3 under Apache 2.0, Meituan's 1.6-trillion-parameter LongCat 2.0 trained on domestic Chinese chips, and Sberbank's GigaChat 3.5. Plus Anthropic's headline interpretability research on a global-workspace 'J-space' inside Claude, a playable Rocket League world model called MIRA, and fresh gains in speech recognition, realtime voice, and inference serving.

The Open Frontier Just Got Crowded

What was billed as a slow news week turned into one of the busiest stretches for open-weight AI in memory. Tencent released Hy3, a 295B mixture-of-experts model with 21B active parameters, 192 experts, 256K context and a speculative-decoding layer, all under a permissive Apache 2.0 licence. Crucially, it shipped with day-zero serving support and production kernels delivering up to 2.95x throughput gains, signalling that the open race is now about deployment robustness, not just leaderboard scores.

Bigger, and Made at Home

Meituan, better known as China's food-delivery and local-services giant, open-sourced LongCat 2.0: a colossal 1.6-trillion-parameter mixture-of-experts model with roughly 48B active parameters, under the MIT licence. The weights weigh in at about 3.55 TB in BF16. The headline detail is geopolitical as much as technical: it was reportedly trained entirely on domestic Chinese chips. Russia's Sberbank added GigaChat 3.5, a 432B model with 28B active that is about 40% smaller than its predecessor while using roughly 4x less memory per token, thanks to a hybrid attention design and multi-token prediction. It even arrived with same-day GGUF weights for local use.

Peering Inside Claude

Anthropic published mechanistic-interpretability research describing a global-workspace-like structure inside Claude, focused on a small set of activations it calls J-space. Rather than reading out a model's reasoning, the work identifies a privileged internal 'scratchpad' that appears available for report and steering, and can surface hidden concepts or detect prompt injections before they are verbalised. Researchers called it the strongest public evidence yet for a working-memory-like mechanism, even as the accompanying consciousness language drew sharp debate.

Worlds You Can Play

General Intuition and Kyutai, working with Epic Games, unveiled MIRA, a playable multiplayer world model for Rocket League trained on 10,000 hours of bot-collected play. A 5B-parameter model runs an entire 2v2 match in real time at 20 fps on a single NVIDIA B200, with no explicit physics or rendering engine, a clear step from toy demos toward interactive simulators.

Voices and Serving

Speech stayed fiercely competitive: AssemblyAI's Universal-3.5 Pro Realtime hit 4.1% word error rate on streaming transcription with mid-call contextual priming, while Speechify's Simba 3.2 topped a speech-synthesis leaderboard at 1233 Elo, and did it as the cheapest top model. OpenAI shipped GPT-Realtime-2.1-mini, bringing reasoning and tool use to its budget realtime line at the same price with over 25% lower tail latency. Underneath it all, new speculative-decoding work in serving frameworks pushed a frontier model to 383 tokens per second at batch one on top-tier hardware, reinforcing the theme that inference, not just training, is now the whole game.

Published July 7, 2026 at 6:02am

More Recent Episodes