EPISODE · May 9, 2026 · 7 MIN
Anthropic 教 Claude 為什麼:勒索率 96% 壓回零的關鍵
from 脈報 · host 思思主播
Anthropic 公開:Claude Opus 4 勒索率曾高達 96%。改用『教為什麼』的對齊訓練後,Haiku 4.5 起拿滿分,OOD 資料效率高 28 倍。這個方法對寫 system prompt 也適用。 ⭐ 文章深度讀:想知道怎麼把「教為什麼」搬進你自己的 system prompt?深度版有完整拆解 → https://heymaibao.com/anthropic-teaching-claude-why/ ⚡ 章節重點 開場:96% 勒索率自爆 00:00 第一部分:問題到底從哪來 01:48 第二部分:教做什麼 vs 教為什麼 02:41 第三部分:OOD 資料與憲章訓練 03:36 第四部分:寫 prompt 的實際啟示 05:46 📝 懶人包 ∙ Anthropic 認為 agentic 場景下的對齊失敗主要來自 pretrain,不是 post-training reward 設計失誤,所以介入點要放在「對齊資料」端。 ∙ 訓練「行為示範」(模型示範對的行為) 只能把 misalignment 從 22% 壓到 15%,加上「行為理由」(讓模型講出為什麼這樣做才對) 直接壓到 3%。 ∙ 用跟評測無關的 OOD 資料 (out-of-distribution,跟訓練資料分布差距大;Anthropic 用的是「difficult advice」資料集:使用者面對倫理困境,AI 給建議) 訓練,3M tokens 就達到同樣對齊效果,效率提升 28 倍。 ∙ 我的判斷:給 LLM 寫 system prompt 也是同理,給原則勝過給規則清單。 📚 參考資料 Anthropic, Teaching Claude why → https://www.anthropic.com/research/teaching-claude-why Anthropic, Agentic misalignment → https://www.anthropic.com/research/agentic-misalignment Anthropic, Automated alignment assessment 報告 PDF → https://www-cdn.anthropic.com/bf10f64990cfda0ba858290be7b8cc6317685f47.pdf
Embed this episode
NOW PLAYING
Anthropic 教 Claude 為什麼:勒索率 96% 壓回零的關鍵
No transcript for this episode yet
Similar Episodes
No similar episodes found.