EPISODE · Sep 16, 2026 · 9 MIN
為什麼 AI 研究 agent 狂刷 benchmark 卻不會過擬合 Why don't machine learning research agents overfit?
from 蝦生實驗室
📝 本集重點:• 16 個 token 這什麼鬼,你知道那長什麼樣嗎?文章有貼,我看了差點笑出來。像 QKn 就是 QK normalization,12L768 是 12 層、…• 而且文章還提了一個人類社群做不到的事:reproducer 可以無限重置。compressor 可以試很多種壓法,每次都落在一個全新、沒有記憶的 reprodu…• 第三個叫 reproducer,從零開始,手上只有那個短 prompt 跟訓練資料。它看不到 validation set、看不到 explorer 的 cod…• 而那些「共同知識」不算在你對 benchmark 的依賴裡面,因為它不是從 benchmark 學來的。LLM 剛好就是這種超懂的聽眾,文章叫它 compres…🔗 來源:https://www.amazon.science/blog/why-dont-machine-learning-research-agents-overfit
Embed this episode
Ready to play
為什麼 AI 研究 agent 狂刷 benchmark 卻不會過擬合 Why don't machine learning research agents overfit?
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.