EPISODE · Sep 15, 2026 · 2 MIN
GPT-Live-1 搭配 Astra medium 以 81.5 分登上 Speech to Speech Index 第 1
from EasyVibeCoding Podcast · host Artificial Analysis
GPT-Live-1 搭配 Astra medium 以 81.5 分登上 Speech to Speech Index 第 1。評估方式 GPT-Live-1 是全雙工語音對語音模型,能在持續對話時把推理與工具使用委派給後端文字模型;開發者透過 API 串流輸入音訊並接收語音回應,後端模型可獨立設定。Artificial Analysis 測試兩種配置:Astra,medium reasoning effort Sol,low reasoning effortSpeech to Speech Index 是四項指標各占 25% 的綜合分數,不是直接平均原始 Elo 與百分比:Big Bench Audio 的 Speech Reasoning、Tau Voice 的 Agentic Performance、Speech Agent Arena 的偏好分數,以及 Task Success Rate。Artificial Analysis 方法說明排行榜表現 GPT-Live-1(Astra, medium)以 81.5 分排名第 1,高於 Grok Voice Think Fast 2.0 High 的 81.3 分;GPT-Live-1(Sol, low)則以 80.1 分排名第 3。分項結果顯示,Astra 在 Tau Voice 的 Agentic Performance 取得 67.9%,Sol 為 59.3%,兩者都高於 Grok Voice Think Fast 2.0 High 的 56.5%,這是 GPT-Live-1 取得整體 Index 領先的重要因素。GPT-Live-1 搭配 Astra medium 後端設定在 Artificial Analysis Speech to Speech Index 以 81.5 分位居第一,領先 Grok Voice Think Fast 2.0 High (81.3 分) 與 GPT-Live-1 (Sol, low) (80.1 分)。在 Big Bench Audio 的音訊推理測試中,Astra 得分 90.1%,Sol 得分 89.0%,低於 Grok Voice Think Fast 2.0 High 的 97.2%與 Qwen Audio 3.0 Realtime Plus 的 99.2%。此外,Full Duplex Bench 子集另行報告,Astra 為 94.9%、Sol 為 97.3%;這些結果不是 Speech to Speech Index 的組成項目。GPT-Live-1(搭配 Astra medium)以 81.5 分位居 Artificial Analysis Speech to Speech Index 榜首,領先 Grok Voice Think Fast 2.0 等模型。對話偏好與任務成功 GPT-Live-1(Sol, low)在 Speech Agent Arena 的對話偏好排名第 3,得分為 1,053 Elo;Astra 的偏好排名第 4,得分為 1,048 Elo。兩者的任務成功率分別為 90.9% 與 87.4%;這是另一項指標,不共用偏好排名。Gemini 3.1 Flash Live Minimal 以 1,096 Elo 領先偏好分數,但 Task Success Rate 為 74.6%;Grok Voice Think Fast 2.0 High 以 94.6%領先任務成功率,偏好分數則為 1,011 Elo。GPT-Live-1 搭配 Astra medium 後端以 81.5 分位居 Speech to Speech Index 第一名;而在圖表展示的 Arena 評測中,GPT-Live-1 (Sol, low) 在 Preference Elo(1053 分)與 Task Success Rate(90.9%)的點估計值均高於 GPT-Live-1 (Astra, medium) 的 1048 分與 87.4%,惟兩者信心區間重疊且各自僅有單次測試數據。Sol 的偏好與任務成功率點估計值都高於 Astra,不過兩者的信賴區間有重疊;因此,這組差異不能脫離目前各配置僅有一次試驗的條件解讀。GPT-Live-1 在 Big Bench Audio 測試集上的首音訊生成時間(Time to First Audio),Sol low 配置為 1.24 秒,Astra medium 配置為 1.34 秒。速度與成本 在 Big Bench Audio 中,GPT-Live-1(Sol, low)平均首次產生音訊需 1.24 秒,Astra 為 1.34 秒;Grok Voice Think Fast 2.0 High 為 0.70 秒,GPT-Realtime-2.1 High 為 1.21 秒。成本則以固定 40 題的 Big Bench Audio 定價子集,依每小時輸入音訊正規化計算:在 Big Bench Audio 子集的每小時輸入音訊成本比較中,Gemini 2.5 Flash Native Audio Dialog 以 $1.42 最低,GPT-Live-1 (Sol, low) 與 GPT-Live-1 (Astra, medium) 分別為 $4.47 與 $5.83,GPT-Realtime-2.1 High, OpenAI 則以 $10.75 最高。Astra:每小時輸入音訊 $5.83 Sol:每小時輸入音訊 $4.47 Grok Voice Think Fast 2.0 High:$4.80 GPT-Realtime-2.1 High:$10.75上述成本包含 GPT-Live-1 的語音工作階段費用,以及依標準費率計算的委派後端文字模型 token 使用量;這是特定基準測試子集的正規化結果,不是一般 API 的每小時費率。在這次測試中,Astra 配置的 Index 與 Tau Voice 分數較高,Sol 則在速度、成本及對話偏好與任務成功率的點估計上較有利,但目前證據仍受單次試驗與重疊信賴區間限制。原文:https://easyvibecoding.app/curated/3340-gpt-live-1-astra-medium-hits-speech-to-speech-index-top
Embed this episode
Ready to play
GPT-Live-1 搭配 Astra medium 以 81.5 分登上 Speech to Speech Index 第 1
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.