Inception Labs 發布 Mercury 2.5,官方稱推理速度達每秒 1,107 tokens episode artwork

EPISODE · Sep 9, 2026 · 1 MIN

Inception Labs 發布 Mercury 2.5,官方稱推理速度達每秒 1,107 tokens

from EasyVibeCoding Podcast · host Inception

Inception Labs 發布 Mercury 2.5,官方稱推理速度達每秒 1,107 tokens。官方公告可參閱〈Introducing Mercury 2.5〉。規格與定價 Mercury 2.5 的主要規格如下:速度:在廣泛可取得的 NVIDIA GPU 上達到 1,107 tokens/sec。 context:260K tokens。 標準價格:每百萬 input tokens $0.20、每百萬 output tokens $0.75。 上線優惠:80% 折扣後為每百萬 input tokens $0.04、每百萬 output tokens $0.15。 能力:可調整 reasoning、平行 tool calls,以及 schema-aligned JSON。公告將 Mercury 2.5 的品質與成本最佳化的 frontier models 相比,包括 GPT-5.6 Luna (Low)、Gemini 3.5 Flash-Lite 與 Claude Haiku 4.5。不過,40% 的提升是對「intelligence」的描述,原文沒有定義單一整體指標或評測方法;1,107 tokens/sec 也只註明使用廣泛可取得的 NVIDIA GPU,未提供硬體型號、服務設定、並行數或測量方法。Mercury 2.5 宣傳短片開場以打字機排列文字點出 LLM 效能瓶頸,隨後展示其具備高推理速度與平行 token 處理能力速度與評測 官方速度圖表列出 Mercury 2.5 為 1,107 tokens/sec,對比 Gemini 3.5 Flash Lite 的 321、Claude Haiku 4.5 的 127,以及 GPT-5.6 Luna (Low) 的 99。與 Mercury 2 的比較圖則顯示:Tau3Bench Telecom:96% 對 65% GPQA Diamond:79% 對 74% IFBench:77% 對 69% AA-LCR:68% 對 41% SciCode:38% 對 37% TerminalBench:37% 對 25% DSQA(10 次 tool calls):34% 對 15% Omniscience Non-Hallucination:33% 對 18% Omniscience Accuracy:22% 對 24% GDPval(Elo):21% 對 13%因此,圖表中的 Mercury 2.5 並非在每一項指標都高於 Mercury 2;Omniscience Accuracy 反而由 24% 降至 22%。Mercury 2.5 在速度 Benchmark 中以每秒 1,107 個 token 領先 Gemini 3.5 Flash Lite(321 tokens/sec)、Claude Haiku 4.5(127 tokens/sec)及 GPT-5.6 Luna (Low)(99 tokens/sec)。Production 案例 Inception Labs 表示,自 Mercury 2 發布後,已有數千名開發者採用、數十家企業投入 production,使用量成長超過一個數量級。這些搜尋、voice 與 coding 工作負載的回饋和 production failure cases,被用來調整 evals 與訓練方向。OpenCall 將 Mercury 用於 AI 電話 Agent;公告稱其 production 工作負載的模型回應中位延遲接近 170 毫秒。OpenCall 另稱,P99 回應時間從數分鐘降至 1 秒,P50 則由 0.4 秒降至低於 0.2 秒。 Augment Code 將 Mercury 用於上下文壓縮(context compaction)、模型路由與 MCP 工具搜尋。改用 Mercury 後,compaction 延遲從約 150 秒降至 27 秒,降低 82%;成本降低 90%,並維持品質,tool-search 摘要則在 1 秒內回傳。OpenCall 與 Augment Code 的數據屬公告引用的客戶報告,提供的資料沒有包含各自的測試設定。Mercury 2.5 在 Tau3Bench Telecom、GPQA Diamond、IFBench、AA-LCR、SciCode、TerminalBench、DSQA (@10 tool calls)、Omniscience Non-Hallucination 與 GDPval (Elo) 基準上領先 Mercury 2,但在 Omniscience Accuracy 以 22% 落後於 Mercury 2 的 24%。新功能與取得方式 Inception Labs 同步預覽 Mercury Voice 與 Mercury Router。Mercury Voice 是針對極低延遲 voice Agent 設計的 dLLM,time-to-first-token(TTFT)低於 170 毫秒;Mercury Router 會理解輸入 prompt,再在 open 與 closed models 之間選擇品質、速度與成本組合最合適的模型。Mercury models 已可透過 Inception API、Baseten 與 OpenRouter 使用;目前提供的資料僅證實 OpenRouter 是取得管道,沒有提供 OpenRouter 的公開 benchmark。企業部署則支援專用容量、自動擴縮、合規控管與可設定的資料保留。Inception Labs 也表示已開始訓練下一個、規模更大的模型,目標在未來數個月發布。Mercury 2.5 宣傳短片開場以打字機排列文字點出 LLM 效能瓶頸,隨後展示其具備高推理速度與平行 token 處理能力 影片中的 Prompt 與操作:Prompt(00:17): 用 HTML-5 產生西洋棋遊戲原文:Generate the game of chess in HTML-5原文:https://easyvibecoding.app/curated/3298-inception-labs-releases-mercury-2-5-hits-1107-tokens-per

Episode metadata supplied by the publisher feed · Published Sep 9, 2026

Embed this episode

Ready to play

Inception Labs 發布 Mercury 2.5,官方稱推理速度達每秒 1,107 tokens

0:00 1:51

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of EasyVibeCoding Podcast?

This episode is 1 minute long.

When was this EasyVibeCoding Podcast episode published?

This episode was published on September 9, 2026.

Can I download this EasyVibeCoding Podcast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!