NVIDIA Releases Nemotron-Labs-3-Puzzle-75B-A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server — 2026-07-09 episode artwork

EPISODE · Jul 9, 2026 · 2 MIN

NVIDIA Releases Nemotron-Labs-3-Puzzle-75B-A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server — 2026-07-09

from Impact Vector: AI Tools · host Alutus LLC

## Short Segments Amazon Science introduces Turnstile, a new tool for reinforcement learning that captures token IDs during agentic interactions. This innovation promises to enhance the precision of RL training by recording exact token-level histories, ensuring models optimize based on accurate past experiences. Coming up, we'll explore Robbyant's latest release in robot manipulation and Datalab's new document extraction tool. Robbyant unveils LingBot-VLA 2.0, an open-source Vision-Language-Action model designed for cross-embodiment robot manipulation. This release aims to bridge the gap between lab success and real-world deployment by enhancing generalization, expanding action spaces, and improving predictive dynamics. With a robust data pipeline and a 6B checkpoint, LingBot-VLA 2.0 is set to advance the capabilities of embodied AI. Datalab introduces Lift, a 9B schema-first extractor that transforms PDFs and images into structured JSON. Unlike traditional document AI tools, Lift focuses on schema-driven extraction, bypassing intermediate representations to deliver application-ready fields directly. This approach positions Lift as a powerful tool for enterprises needing precise data extraction from complex documents. ## Feature Story NVIDIA's release of Nemotron-Labs-3-Puzzle-75B-A9B marks a significant leap in large hybrid MoE model efficiency. This compressed variant of the Nemotron-3-Super model achieves over double the server throughput while maintaining user throughput, thanks to a reduction in active parameters from 12.8B to 9.3B. The model's architecture preserves the original's 88-block layout, optimizing capacity within these blocks to enhance performance. The development targets two key performance metrics: doubling server throughput at 100 tokens per second per user and supporting eight concurrent 1M-token requests on a single H100. This is achieved through a strategic reduction in model weight from 70 GB to 44.5 GB, allowing for increased concurrency and efficiency. The iterative Puzzle approach used in this compression process outperforms single-step methods, offering a 0.57-point improvement at the same compression target. For developers and enterprises, this means more efficient deployment of large-scale models without sacrificing quality. The ability to handle more users concurrently at a lower computational cost could transform how AI services are delivered, making them more accessible and scalable. As NVIDIA continues to refine these models, the focus will likely remain on balancing performance with resource efficiency, a critical factor for widespread AI adoption.

Episode metadata supplied by the publisher feed · Published Jul 9, 2026

Embed this episode

Ready to play

NVIDIA Releases Nemotron-Labs-3-Puzzle-75B-A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server — 2026-07-09

0:00 2:57

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Impact Vector: AI Tools?

This episode is 2 minutes long.

When was this Impact Vector: AI Tools episode published?

This episode was published on July 9, 2026.

Can I download this Impact Vector: AI Tools episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!