EPISODE · Aug 26, 2026 · 10 MIN
Z.ai 發布 GLM-5.3-Flash:以 18B 啟用參數支援原生多模態與 1M-token 長上下文
from EasyVibeCoding Podcast · host Z.ai
Z.ai 發布 GLM-5.3-Flash:以 18B 啟用參數支援原生多模態與 1M-token 長上下文。 Z.ai 在 2026 年 8 月 26 日的貼文中強調,這款模型以較低成本提供 coding 與長跨度 Agent 任務能力,並完全在中國 AI 晶片上執行。 正式發布與定位 GLM-5.3-Flash 是 GLM-5 系列首個原生 multimodal model,採用 320B-A18B 架構,也就是總參數 320B、啟用參數 18B,並以 MIT License 發布。它支援 1M-token context window,最大輸出長度為 131K token,預設使用 max reasoning。Z.ai 表示,模型先前曾以 Ox Alpha 名稱預覽,且整個執行過程都由 Chinese AI chips 提供支援。 模型已可在多個官方平台使用: Weights:Hugging Face 模型頁面 API:GLM-5.3-Flash API 文件 Coding Plan:訂閱頁面 ZCode:ZCode Chat:Z.ai Chat AutoClaw:AutoClaw 產品文章:Z.ai 官方 blog 技術報告:Technical report 引用資料另以 2026 年 2 月 17 日記載發布時間,與 Z.ai 貼文標示的 2026 年 8 月 26 日不一致;但兩者都指向同一款 GLM-5.3-Flash,以及它由 ox-alpha 測試版本轉為正式公開版本的脈絡。 價格與使用管道 Z.ai 貼出的 Standard API Pricing 以每 1M tokens 計算: Input:$0.15 Output:$0.50 Cached input:$0.03 OpenRouter 的公告則指出,發布優惠至 Sep 9 at 16:00 UTC 前為五折: Input:$0.075/M Output:$0.25/M Cached input:$0.015/M 優惠期結束後,OpenRouter 預告回到 $0.15/M、$0.50/M 與 $0.03/M。Zixuan Li 的貼文也表示,官方 Z.ai API 在接下來兩週提供 50% 折扣,第三方 model aggregators 同樣可使用優惠;不過該貼文在「After the discount」標籤下列出的仍是 $0.075、$0.25 與 $0.015,與 Z.ai 及 OpenRouter 公告的標準價格描述不一致,使用者應以實際帳單與方案頁面為準。 來源:@ZixuanLi(回覆)|GLM-5.3-Flash 在 Input($0.15)、Output($0.50)與 Cached Input($0.03)價格上皆低於 GLM-5.2、Claude Sonnet 5、GPT-5.6 Terra 與 Gemini 3.7 Flash。 模型也已加入 OpenCode Go,OpenCode 表示限時提供 double usage。這項額外用量屬於 OpenCode Go 的限時方案,不等同於所有平台都提供相同配額。 模型架構與效率 GLM-5.3-Flash 從重新訓練的 base model 開始,重新設計 architecture 與 training recipe。核心變化是首次在 GLM 系列採用 sparse attention 與 linear attention 組成的 hybrid architecture:linear attention 以 state modeling 捕捉 local dependencies,sparse attention 則透過 lightweight indexer 擷取 global context,藉此降低長上下文服務的計算成本,同時維持 long-context 能力。模型也導入 Manifold-Constrained Hyper-Connections(mHC),並使用最新的 30T-token multimodal pre-training corpus。 相較 GLM-5.3,Z.ai 文件列出的效率改善包括: Attention computation 降低 3.01×,其他文件約寫為 3.0×。 KV cache size 降低 4.44×,其他文件約寫為 4.4×。 在 1M-token context 下,IndexPool 會以 weighted pooling 將 4 個 indexer key vectors 壓縮成 1 個。 GLM-5.3-Flash 架構在長序列下顯著降低每層 KV-Cache 大小與 Attention 計算負擔,表現優於 GLM-5.3。 與 GLM-4.5 series 相比,總參數量為 320B 對 355B,啟用參數則由 32B 降至 18B,layers 也由 92 層降至 45 層。來源指出,GLM-5.3-Flash 的 attention compute 是比較模型中最低,但 KV cache 仍略大於 Kimi-K3 與 DeepSeek-V4-Flash,因此並非所有效率指標都已達到最佳。 可支援本地部署的 frameworks 包括 SGLang、vLLM、TokenSpeed 與 KTransformers。Z.ai 也表示,這是首個採用 sparse attention 與 linear attention hybrid architecture 的 open-source frontier model。 API 與多模態能力 API model code 為 glm-5.3-flash。除了文字,它原生支援圖片、影片與檔案;圖片可透過 messages[].content[] 的 type: imageurl content block 傳入,imageurl.url 可接受 image URL 或 Base64 Data URL,也能放入多個 image blocks。官方建議設定如下: `json { "temperature": 1, "top_p": 0.95, "reasoning_effort": "max", "thinking.type": "enabled", "thinking.clear_thinking": false, "stream": true, "tool_stream": true } ` 其中 thinking.type 只支援 enabled,thinking 不能停用;若使用 streaming requests,官方建議同時啟用 stream: true 與 tool_stream: true。相關介面與能力文件包括: Chat Completion API Thinking Mode Streaming Output Function Calling Context Caching Structured Output 完整文件索引:llms.txt 在 GLM Coding Plan 中,模型已 fully available,quota 為 GLM-5.3 的 3 倍;off-peak hours,包括週末全天,points-based quota 只消耗標準 points 的 50%。可參考 Personal Plan 與 Team Plan。 Coding 與 benchmark 表現 Z.ai 宣稱 GLM-5.3-Flash 在 coding 與 agentic benchmarks 多數勝過 GLM-5.2,且接近 Claude Opus 4.8。Artificial Analysis Intelligence Index v4.1.1 的得分為 57,每項 task 的 discounted 價格為 $0.045;來源據此估計,同等 intelligence 過去約需 10 倍成本。 具體結果包括: DeepSWE v1.1:63.4,相較 GLM-5.2 的 46.2。 AutomationBench:48.8,相較 GLM-5.2 的 26.2。 Z.ai Code Bench v1.0 在 Claude Code 2.1.207 上執行時,各 effort level 均勝過 GLM-5.2;max effort 為 29.0,Claude Opus 4.8 為 29.5。 Z.ai Code Bench 被 Z.ai 描述為衡量 real-world coding performance 的評測,官方稱 GLM-5.3-Flash 在每個 effort level 都優於 GLM-5.2,…
Embed this episode
Ready to play
Z.ai 發布 GLM-5.3-Flash:以 18B 啟用參數支援原生多模態與 1M-token 長上下文
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.