You're offline - Playing from downloaded podcasts
Back to All Episodes
Podcast Episode

Meta's Muse Goes Agentic, NVIDIA's Audex Audio Brain, and the AI That Checks Its Own Homework

July 9, 2026

0:00
10:18
Podcast Thumbnail

Meta Superintelligence Labs launches Muse Image and previews Muse Video with a self-refining, agentic generation loop, while NVIDIA ships Audex, a 30B mixture-of-experts audio model, and a Puzzle compression method that doubles server throughput. We also cover Cohere's open-source Arabic speech recognition, Liquid AI's clever 'Antidoom' fix for reasoning loops, open robotics moving into shared platforms, a humbling new legal benchmark, and the rise of background agents from Anthropic and Google.

Meta's Muse Turns Image Generation Into an Agent

Meta Superintelligence Labs launched Muse Image and previewed Muse Video, and the headline isn't picture quality, it's the process. Muse plans, searches the web, uses tools, runs code, and refines its own output before rendering, behaviour that emerged during reinforcement learning rather than being hand-scripted. Muse Image hit #2 on Image Arena behind GPT Image 2, with Muse Video debuting at #3 on Video Arena.

NVIDIA Audex: One Backbone for Text and Audio

NVIDIA released Audex, a 30B-parameter mixture-of-experts model with just 3B active parameters and a 1M-token context. Its core claim is preserving text intelligence while adding broad audio generation and understanding on a single unified backbone, avoiding the usual trade-off where bolting on sound dents a model's language skills.

Cohere Opens Up Arabic Speech Recognition

Cohere launched Cohere Transcribe Arabic under an Apache 2.0 licence, billed as the most accurate open-source Arabic speech recognition model. It targets the language's many dialects, code-switching, and Arabic-accented English, putting capable voice tech in local hands for hundreds of millions of underserved speakers.

Liquid AI's Antidoom Kills the Reasoning Stutter

Liquid AI shipped Antidoom, an open-source training method that stops small reasoning models entering 'doom loops' where they repeat tokens until context runs out. The technique, Final Token Preference Optimization, relabels the loop-triggering token and redistributes probability. Reported looping dropped from around 10.2% to 1.4% on one model and from roughly 22.9% to 1% on another.

NVIDIA Puzzle: Compression Without the Brain Damage

NVIDIA's Puzzle compression squeezed a ~120B model down to ~75B while retaining reasoning, coding, long-context, and agentic quality. The payoff is roughly 2x server throughput, with single-chip concurrency at 1M-token context rising from 1 request to 8.

Open Robotics Consolidates

NVIDIA brought its GR00T robot foundation model and Isaac Teleop tools into an open robotics platform, part of a broader consolidation around shared tools and models. One team showcased a full robot prototype built by a small group in just 9 months, evidence that open plumbing is lowering the barrier to building capable machines.

A Reality Check From Legal AI

A new legal-agent benchmark tested models across 120 private legal tasks spanning 24 practice areas. Even the strongest performer fully passed only around 14.2% of tasks end-to-end. The lesson: models can satisfy many individual rubric items yet still fail to deliver an acceptable finished deliverable.

The Rise of Background Agents

Anthropic brought Claude Cowork to mobile and web, reframing the assistant as a background teammate rather than a foreground chat. Google pushed the same direction, adding background execution and other agent infrastructure to its Gemini API. The emerging view: much of the real capability now lives in the harness around the model, not just the raw weights.

Published July 9, 2026 at 1:38am

More Recent Episodes