GPT-6 Astra 在 ARC-AGI-3 得分 99.9%,與 Sol 採不同評測框架 episode artwork

EPISODE · Sep 5, 2026 · 1 MIN

GPT-6 Astra 在 ARC-AGI-3 得分 99.9%,與 Sol 採不同評測框架

from EasyVibeCoding Podcast · host Tibo

GPT-6 Astra 在 ARC-AGI-3 得分 99.9%,與 Sol 採不同評測框架。圖表數據 2026 年 9 月 3 日,Tibo 分享的比較圖標示 GPT‑6 Astra 99.9%、Claude Opus 5 30.2%、GPT‑5.6 Sol 7.8%;圖說另列平均人類測試者 48%。OpenAI 官方發布頁也列出 Astra 在 ARC-AGI-3 得分 99.9%,並寫到它在 96% 的關卡超過人類行動效率基準。GPT-6 Astra 在 ARC-AGI-3 得分 99.9%;圖中另列 Claude Opus 5(30.2%)與 GPT-5.6 Sol(7.8%)。Astra 使用 Responses API 執行框架,與圖中 Sol 的測法不同;OpenAI 估計 Sol 若採同一框架,得分約為 30%。評測脈絡 ARC-AGI-3 測試模型學習陌生的互動任務;Astra 的 99.9% 是透過 Responses API 執行框架測得。圖說另提供同一執行框架下 GPT‑5.6 Sol 約 30% 的估計,因此圖表的 7.8% 是不同脈絡下的另一個數值,不能和約 30% 混用。原始資料沒有 ARC-AGI-3 的執行設定、樣本數或信賴區間,圖表的方法與排名也尚未獲獨立驗證。解讀方式 Tibo 問「我們需要另一個 AGI 評測,下一個門檻又會移到哪裡?」這是對評測目標是否移動的評論,不是數字的額外證明。圖表只能說明 Astra 在指定任務和執行框架下的分數;比較時應同時記錄模型分數、人類基準、Responses API 執行框架、任務設計與計分規則,不能據此推論模型已具備可泛化的 AGI 能力。原文:https://easyvibecoding.app/curated/3269-gpt-6-astra-hits-99-9-arc-agi-3-responses-api-harness

Episode metadata supplied by the publisher feed · Published Sep 5, 2026

Embed this episode

Ready to play

GPT-6 Astra 在 ARC-AGI-3 得分 99.9%,與 Sol 採不同評測框架

0:00 1:14

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of EasyVibeCoding Podcast?

This episode is 1 minute long.

When was this EasyVibeCoding Podcast episode published?

This episode was published on September 5, 2026.

Can I download this EasyVibeCoding Podcast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!