EPISODE · Sep 9, 2026 · 1 MIN
Inception Labs 發布 Mercury 2.5,官方稱推理速度達每秒 1,107 tokens
from EasyVibeCoding Podcast · host Inception
Inception Labs 發布 Mercury 2.5,官方稱推理速度達每秒 1,107 tokens。官方公告可參閱〈Introducing Mercury 2.5〉。規格與定價 Mercury 2.5 的主要規格如下:速度:在廣泛可取得的 NVIDIA GPU 上達到 1,107 tokens/sec。 context:260K tokens。 標準價格:每百萬 input tokens $0.20、每百萬 output tokens $0.75。 上線優惠:80% 折扣後為每百萬 input tokens $0.04、每百萬 output tokens $0.15。 能力:可調整 reasoning、平行 tool calls,以及 schema-aligned JSON。公告將 Mercury 2.5 的品質與成本最佳化的 frontier models 相比,包括 GPT-5.6 Luna (Low)、Gemini 3.5 Flash-Lite 與 Claude Haiku 4.5。不過,40% 的提升是對「intelligence」的描述,原文沒有定義單一整體指標或評測方法;1,107 tokens/sec 也只註明使用廣泛可取得的 NVIDIA GPU,未提供硬體型號、服務設定、並行數或測量方法。Mercury 2.5 宣傳短片開場以打字機排列文字點出 LLM 效能瓶頸,隨後展示其具備高推理速度與平行 token 處理能力速度與評測 官方速度圖表列出 Mercury 2.5 為 1,107 tokens/sec,對比 Gemini 3.5 Flash Lite 的 321、Claude Haiku 4.5 的 127,以及 GPT-5.6 Luna (Low) 的 99。與 Mercury 2 的比較圖則顯示:Tau3Bench Telecom:96% 對 65% GPQA Diamond:79% 對 74% IFBench:77% 對 69% AA-LCR:68% 對 41% SciCode:38% 對 37% TerminalBench:37% 對 25% DSQA(10 次 tool calls):34% 對 15% Omniscience Non-Hallucination:33% 對 18% Omniscience Accuracy:22% 對 24% GDPval(Elo):21% 對 13%因此,圖表中的 Mercury 2.5 並非在每一項指標都高於 Mercury 2;Omniscience Accuracy 反而由 24% 降至 22%。Mercury 2.5 在速度 Benchmark 中以每秒 1,107 個 token 領先 Gemini 3.5 Flash Lite(321 tokens/sec)、Claude Haiku 4.5(127 tokens/sec)及 GPT-5.6 Luna (Low)(99 tokens/sec)。Production 案例 Inception Labs 表示,自 Mercury 2 發布後,已有數千名開發者採用、數十家企業投入 production,使用量成長超過一個數量級。這些搜尋、voice 與 coding 工作負載的回饋和 production failure cases,被用來調整 evals 與訓練方向。OpenCall 將 Mercury 用於 AI 電話 Agent;公告稱其 production 工作負載的模型回應中位延遲接近 170 毫秒。OpenCall 另稱,P99 回應時間從數分鐘降至 1 秒,P50 則由 0.4 秒降至低於 0.2 秒。 Augment Code 將 Mercury 用於上下文壓縮(context compaction)、模型路由與 MCP 工具搜尋。改用 Mercury 後,compaction 延遲從約 150 秒降至 27 秒,降低 82%;成本降低 90%,並維持品質,tool-search 摘要則在 1 秒內回傳。OpenCall 與 Augment Code 的數據屬公告引用的客戶報告,提供的資料沒有包含各自的測試設定。Mercury 2.5 在 Tau3Bench Telecom、GPQA Diamond、IFBench、AA-LCR、SciCode、TerminalBench、DSQA (@10 tool calls)、Omniscience Non-Hallucination 與 GDPval (Elo) 基準上領先 Mercury 2,但在 Omniscience Accuracy 以 22% 落後於 Mercury 2 的 24%。新功能與取得方式 Inception Labs 同步預覽 Mercury Voice 與 Mercury Router。Mercury Voice 是針對極低延遲 voice Agent 設計的 dLLM,time-to-first-token(TTFT)低於 170 毫秒;Mercury Router 會理解輸入 prompt,再在 open 與 closed models 之間選擇品質、速度與成本組合最合適的模型。Mercury models 已可透過 Inception API、Baseten 與 OpenRouter 使用;目前提供的資料僅證實 OpenRouter 是取得管道,沒有提供 OpenRouter 的公開 benchmark。企業部署則支援專用容量、自動擴縮、合規控管與可設定的資料保留。Inception Labs 也表示已開始訓練下一個、規模更大的模型,目標在未來數個月發布。Mercury 2.5 宣傳短片開場以打字機排列文字點出 LLM 效能瓶頸,隨後展示其具備高推理速度與平行 token 處理能力 影片中的 Prompt 與操作:Prompt(00:17): 用 HTML-5 產生西洋棋遊戲原文:Generate the game of chess in HTML-5原文:https://easyvibecoding.app/curated/3298-inception-labs-releases-mercury-2-5-hits-1107-tokens-per
Embed this episode
Ready to play
Inception Labs 發布 Mercury 2.5,官方稱推理速度達每秒 1,107 tokens
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.