EPISODE · Jul 8, 2026 · 3 MIN
NVIDIA Releases Audex (Nemotron-Labs-Audex-30B-A3B): A Unified Audio-Text LLM That Preserves the Text — 2026-07-08
from Impact Vector: AI Tools · host Alutus LLC
## Short Segments Ant Group's Robbyant has open-sourced LingBot-Vision, a vision foundation model that prioritizes boundary-centric perception. Unlike traditional models that focus on semantic invariance, LingBot-Vision emphasizes fine-grained spatial structures, crucial for robots and embodied systems. This 1B-parameter model, available on Hugging Face, matches or surpasses models up to seven times larger on dense spatial tasks. By treating boundaries as a native pretraining signal, it offers a new approach to spatial perception, potentially transforming how robots interpret their environments. NVIDIA's Cosmos-Framework tutorial offers a Colab-friendly approach to understanding Cosmos 3 world models. While full Cosmos 3 inference isn't feasible on standard Colab hardware, the tutorial provides a hands-on miniature implementation using the framework's structure and model modes. This approach allows users to build and train a compact omnimodal Mixture-of-Transformers world model, demonstrating cross-modal attention and expert routing for text, vision, and action streams. It's a practical way to explore the core ideas of Cosmos 3 without needing high-end hardware. ## Feature Story NVIDIA has unveiled Audex, a unified audio-text large language model that maintains the text intelligence of its backbone while integrating audio capabilities. This release addresses a common challenge in multimodal models, where adding audio or vision outputs often leads to a drop in text benchmark performance. Audex, however, is designed to avoid this regression, offering a model that handles both audio and text without compromising on text intelligence. Audex is a 30B-parameter Mixture-of-Experts model that processes audio inputs by encoding them into the text embedding space, treating audio outputs as text tokens. This approach ensures that text scores remain consistent with the backbone, with only minor variations across benchmarks. The model employs a multi-stage supervised fine-tuning process and text-only Cascade Reinforcement Learning to maintain its performance across modalities. What sets Audex apart is its ability to generate general audio beyond speech, making it one of the few open models with this capability. By integrating audio understanding, speech recognition, translation, text-to-speech, and audio generation, Audex offers a comprehensive solution for developers and enterprises looking to leverage multimodal AI. This development is particularly significant for industries that require seamless integration of audio and text processing, such as media, entertainment, and customer service. As NVIDIA continues to push the boundaries of AI with models like Audex, the potential for more efficient and accurate multimodal systems becomes increasingly tangible. For developers, this means access to a powerful tool that can enhance applications with advanced audio and text capabilities, all while maintaining high performance standards. Looking ahead, the release of Audex under a noncommercial license opens up opportunities for further research and innovation in the field of multimodal AI. Stay tuned as we continue to track the impact of this groundbreaking model on the AI landscape.
Embed this episode
Ready to play
NVIDIA Releases Audex (Nemotron-Labs-Audex-30B-A3B): A Unified Audio-Text LLM That Preserves the Text — 2026-07-08
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.