@Mnilax:Karpathy 的 4 條 CLAUDE.md 規則將 Claude 的錯誤率從 41% 降至 11%。在測試 30 個程式庫後,我新增了 8 條
 
 20… episode artwork

EPISODE · May 9, 2026 · 7 MIN

@Mnilax:Karpathy 的 4 條 CLAUDE.md 規則將 Claude 的錯誤率從 41% 降至 11%。在測試 30 個程式庫後,我新增了 8 條 20…

from EasyVibeCoding Podcast · host Mnimiy

Karpathy 的 4 條 CLAUDE.md 規則將 Claude 的錯誤率從 41% 降至 11%。在測試 30 個程式庫後,我新增了 8 條 2026 年 1 月下旬,Andrej Karpathy 發布了一則討論串,抱怨 Claude 撰寫程式碼的方式。他指出了三種失敗模式:默認錯誤的假設、過度複雜化,以及對不該更動的程式碼造成了連帶損害。 Forrest Chang 讀了這則討論串,將這些抱怨歸納為 4 條行為準則,寫入一個 CLAUDE.md 文件中,並發布到 GitHub 上。它在第一天就獲得了 5,828 個星號,兩週內累積 60,000 個書籤,至今已有 120,000 個星號。這是 2026 年成長最快的單文件儲存庫。 接著,我在 6 週內於 30 個程式庫中測試了這套規則。 這 4 條規則確實有效。在發揮其優勢的任務中,原本約 40% 的錯誤率下降到 3% 以下。但該模板是為了修正 2026 年 1 月的程式碼撰寫錯誤而設計的。 2026 年 5 月的 Claude Code 生態系統面臨不同的問題——Agent 衝突、Hook 連鎖反應、技能載入衝突,以及在不同對話階段中斷的多步驟工作流程。 因此,我新增了 8 條規則。以下是完整的 12 條規則 CLAUDE.md,說明每一條規則為何存在,以及原始 Karpathy 模板在哪些地方會默默失效。 如果你想跳過解釋直接複製,完整文件在文末。 為什麼這很重要 Claude Code 的 CLAUDE.md 是整個 AI 程式開發堆疊中最被低估的文件。大多數開發者要麼: 將其視為所有偏好的垃圾桶,膨脹到 4,000 個 token 以上,導致合規性降至 30%。 完全忽略它,每次都重新下 Prompt——浪費 5 倍的 token,且對話階段之間缺乏一致性。 複製一次模板後就拋諸腦後。這在兩週內有效,但隨著程式庫變動,它會默默失效。 Anthropic 官方文件明確指出:CLAUDE.md 僅供參考。Claude 大約有 80% 的時間會遵守它。一旦超過 200 行,合規性會急劇下降,因為重要的規則被淹沒在雜訊中。 Karpathy 的模板用一個 65 行、4 條規則的文件解決了這個問題。這是基準線。 上限可以更高。透過我下面介紹的額外 8 條規則,你不僅能涵蓋 Karpathy 當初抱怨的 2026 年 1 月程式碼撰寫問題,還能解決模板撰寫時尚未出現的 2026 年 5 月 Agent 編排問題。 原始的 4 條規則 如果你還沒讀過 Forrest Chang 的儲存庫,這是基準: 規則 1 — 程式撰寫前先思考。 沒有默認假設。說明你的假設。提出權衡考量。在猜測前先詢問。當存在更簡單的方法時,提出反對意見。 規則 2 — 簡潔優先。 用最少的程式碼解決問題。不要有預測性的功能。不要為單次使用的程式碼建立抽象層。如果資深工程師會覺得它過於複雜——那就簡化它。 規則 3 — 外科手術式的變更。 只更動必須更動的部分。不要「優化」相鄰的程式碼、註解或格式。不要重構沒壞的東西。符合現有的風格。 規則 4 — 目標導向執行。 定義成功標準。循環直到驗證完成。不要告訴 Claude 該遵循什麼步驟,告訴它成功是什麼樣子,讓它自行迭代。 這四條規則解決了我觀察到的非監督式 Claude Code 對話中約 40% 的失敗模式。剩下的約 60% 存在於以下的空白地帶。 我新增的 8 條規則(以及原因) 每一條都源自 Karpathy 的 4 條規則不足以應付的真實時刻。我將展示該場景,然後給出規則。 規則 5 — 不要讓模型做非語言類的工作 Karpathy 的規則對此隻字未提。模型會決定那些應該由確定性程式碼處理的事,例如是否重試 API 呼叫、如何路由訊息、何時升級處理。每週的決定都不一樣。在每個 token 0.003 美元的成本下,這種 if-else 邏輯很不穩定。 ` Rule 5 — Use the model only for judgment calls Use Claude for: classification, drafting, summarization, extraction from unstructured text. Do NOT use Claude for: routing, retries, status-code handling, deterministic transforms. If a status code already answers the question, plain code answers the question. ` 場景:一段呼叫 Claude 來「決定是否在 503 錯誤時重試」的程式碼,前兩週運作良好,後來開始不穩定,因為模型開始讀取請求主體作為決策的 context。重試策略變得隨機,因為 Prompt 是隨機的。 規則 6 — 硬性 token 預算,沒有例外 沒有預算的 CLAUDE.md 就像一張空白支票。每個循環都有可能演變成 50,000 個 token 的 context 傾倒。模型不會自己停下來。 ` Rule 6 — Token budgets are not advisory Per-task budget: 4,000 tokens. Per-session budget: 30,000 tokens. If a task is approaching budget, summarize and start fresh. Do not push through. Surfacing the breach > silently overrunning. ` 場景:一個除錯對話持續了 90 分鐘。模型很樂意在同一個 8KB 的錯誤訊息上不斷迭代,逐漸忘記它已經嘗試過哪些修復方法。到最後,它建議的修復方案是我 40 條訊息前就已經拒絕過的。token 預算本可以在第 12 分鐘就終止它。 規則 7 — 呈現衝突,不要折衷 當程式庫的兩個部分意見不合時,Claude 會試圖討好兩者。結果就是前後不一致。 ` Rule 7 — Surface conflicts, don't average them If two existing patterns in the codebase contradict, don't blend them. Pick one (the more recent / more tested), explain why, and flag the other for cleanup. "Average" code that satisfies both rules is the worst code. ` 場景:一個程式庫有兩種錯誤處理模式——一種是帶有明確 try/catch 的 async/await,另一種是全域錯誤邊界。Claude 寫出的新程式碼兩者都用了。錯誤處理器重複了。我花了 30 分鐘才弄清楚為什麼錯誤被吞掉了兩次。 規則 8 — 撰寫前先閱讀 Karpathy 的「外科手術式的變更」告訴 Claude 不要更動相鄰程式碼。但它沒告訴 Claude 要先理解相鄰程式碼。沒有這一條,Claude 寫出的新程式碼會與 30 行外的現有程式碼衝突。 ` Rule 8 — Read before you write Before adding code in a file, read the file's exports, the immediate caller, and any obvious shared utilities. If you don't understand why existing code is structured the way it is, ask before adding to it. "Looks orthogonal to me" is the most dangerous phrase in this codebase. ` 場景:Claude 在一個它沒讀過的現有函數旁邊新增了一個一模一樣的函數。兩個函數做同樣的事。因為匯入順序的關係,新的函數優先被執行。而舊的函數已經是 6 個月來的唯一真理。 規則 9 — 測試不是選配,但也不是最終目標 Karpathy 的「目標導向執行」暗示測試是成功標準。實際上,Claude 把「測試通過」當成唯一目標,寫出的程式碼能通過淺層測試,卻破壞了其他所有東西。 ` Rule 9 — Tests verify intent, not just behavior Every…

Episode metadata supplied by the publisher feed · Published May 9, 2026

Embed this episode

Ready to play

@Mnilax:Karpathy 的 4 條 CLAUDE.md 規則將 Claude 的錯誤率從 41% 降至 11%。在測試 30 個程式庫後,我新增了 8 條 20…

0:00 7:19

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of EasyVibeCoding Podcast?

This episode is 7 minutes long.

When was this EasyVibeCoding Podcast episode published?

This episode was published on May 9, 2026.

Is there a transcript available for this episode?

Yes, a full transcript is available for this episode. You can read the complete transcript on the episode page.

Can I download this EasyVibeCoding Podcast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!