@NVIDIAAI:NVIDIA 推出 Nemotron-Labs-TwoTower 模型,啟用雙塔平行生成,達到 2.42 倍推論速度。
 
 核心架構與運作機制
 NVIDIA… episode artwork

EPISODE · Jul 1, 2026 · 2 MIN

@NVIDIAAI:NVIDIA 推出 Nemotron-Labs-TwoTower 模型,啟用雙塔平行生成,達到 2.42 倍推論速度。 核心架構與運作機制 NVIDIA…

from EasyVibeCoding Podcast · host NVIDIA AI

NVIDIA 推出 Nemotron-Labs-TwoTower 模型,啟用雙塔平行生成,達到 2.42 倍推論速度。 核心架構與運作機制 NVIDIA Research 發布的「Nemotron-Labs-TwoTower」模型,是基於「Nemotron-3-Nano-30B-A3B」骨幹所開發的區塊式自回歸擴散語言模型。該模型將傳統單一模型的任務拆分為兩個獨立的塔: Context Tower(AR/Context):保持凍結狀態,負責處理乾淨的 Prompt 與已提交的 token,並提供 KV 快取與 Mamba-2 狀態。 Denoiser Tower(Diffusion/Denoiser):為可訓練塔,透過遮罩擴散(mask diffusion)機制,以區塊為單位進行平行去噪與生成。 效能表現與技術細節 根據 NVIDIA 的測試數據,該模型在維持原始自回歸基準模型 98.7% 品質的前提下,達成了 2.42 倍的實際生成吞吐量。其運作流程如下: 這張圖表展示了結合 AR Context Tower 與 Diffusion / Denoiser Tower 的運作流程,包含 AR Prefill、Diffusion Denoising 以及 AR Context Update 三個階段。 AR Prefill:Context Tower 先對輸入序列進行編碼,並在每一層輸出注意力 KV 快取與 Mamba-2 邊界狀態。 Diffusion Denoising:Denoiser Tower 針對區塊內的 [MASK] token 進行多次迭代去噪。過程中透過雙向區塊內注意力(bidirectional in-block attention)、層對齊的交叉注意力(cross-attention)與 Context Tower 互動,並由 Context Tower 的狀態進行種子初始化。 AR Context Update:當區塊生成完成並確認後,將結果提交並更新 Context Tower 的 KV 與 Mamba 快取,隨後進入下一個區塊的處理。 模型取得與應用 此模型已於 Hugging Face 開放下載,並附帶相關研究論文 arXiv 2606.26493。該專案採用 NVIDIA Nemotron 開放模型授權協議,適用於商業用途,用來為開發專業 AI Agent 提供更具效率的解決方案。這張圖表展示了結合 AR Context Tower 與 Diffusion / Denoiser Tower 的運作流程,包含 AR Prefill、Diffusion Denoising 以及 AR Context Update 三個階段。 影片中的 Prompt 與操作:操作步驟: 1. (00:00)AR Prefill 階段處理輸入序列。 2. (00:03)Diffusion Denoising 階段處理 [M] 標記並與 AR Tower 交互。 3. (00:09)AR Context Update 階段更新上下文資訊。原文:https://easyvibecoding.app/curated/2291

Episode metadata supplied by the publisher feed · Published Jul 1, 2026

Embed this episode

Ready to play

@NVIDIAAI:NVIDIA 推出 Nemotron-Labs-TwoTower 模型,啟用雙塔平行生成,達到 2.42 倍推論速度。 核心架構與運作機制 NVIDIA…

0:00 2:04

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of EasyVibeCoding Podcast?

This episode is 2 minutes long.

When was this EasyVibeCoding Podcast episode published?

This episode was published on July 1, 2026.

Can I download this EasyVibeCoding Podcast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!