Z.ai 發布 GLM-5.3-Flash:以 18B 啟用參數支援原生多模態與 1M-token 長上下文 episode artwork

EPISODE · Aug 26, 2026 · 10 MIN

Z.ai 發布 GLM-5.3-Flash:以 18B 啟用參數支援原生多模態與 1M-token 長上下文

from EasyVibeCoding Podcast · host Z.ai

Z.ai 發布 GLM-5.3-Flash:以 18B 啟用參數支援原生多模態與 1M-token 長上下文。 Z.ai 在 2026 年 8 月 26 日的貼文中強調,這款模型以較低成本提供 coding 與長跨度 Agent 任務能力,並完全在中國 AI 晶片上執行。 正式發布與定位 GLM-5.3-Flash 是 GLM-5 系列首個原生 multimodal model,採用 320B-A18B 架構,也就是總參數 320B、啟用參數 18B,並以 MIT License 發布。它支援 1M-token context window,最大輸出長度為 131K token,預設使用 max reasoning。Z.ai 表示,模型先前曾以 Ox Alpha 名稱預覽,且整個執行過程都由 Chinese AI chips 提供支援。 模型已可在多個官方平台使用: Weights:Hugging Face 模型頁面 API:GLM-5.3-Flash API 文件 Coding Plan:訂閱頁面 ZCode:ZCode Chat:Z.ai Chat AutoClaw:AutoClaw 產品文章:Z.ai 官方 blog 技術報告:Technical report 引用資料另以 2026 年 2 月 17 日記載發布時間,與 Z.ai 貼文標示的 2026 年 8 月 26 日不一致;但兩者都指向同一款 GLM-5.3-Flash,以及它由 ox-alpha 測試版本轉為正式公開版本的脈絡。 價格與使用管道 Z.ai 貼出的 Standard API Pricing 以每 1M tokens 計算: Input:$0.15 Output:$0.50 Cached input:$0.03 OpenRouter 的公告則指出,發布優惠至 Sep 9 at 16:00 UTC 前為五折: Input:$0.075/M Output:$0.25/M Cached input:$0.015/M 優惠期結束後,OpenRouter 預告回到 $0.15/M、$0.50/M 與 $0.03/M。Zixuan Li 的貼文也表示,官方 Z.ai API 在接下來兩週提供 50% 折扣,第三方 model aggregators 同樣可使用優惠;不過該貼文在「After the discount」標籤下列出的仍是 $0.075、$0.25 與 $0.015,與 Z.ai 及 OpenRouter 公告的標準價格描述不一致,使用者應以實際帳單與方案頁面為準。 來源:@ZixuanLi(回覆)|GLM-5.3-Flash 在 Input($0.15)、Output($0.50)與 Cached Input($0.03)價格上皆低於 GLM-5.2、Claude Sonnet 5、GPT-5.6 Terra 與 Gemini 3.7 Flash。 模型也已加入 OpenCode Go,OpenCode 表示限時提供 double usage。這項額外用量屬於 OpenCode Go 的限時方案,不等同於所有平台都提供相同配額。 模型架構與效率 GLM-5.3-Flash 從重新訓練的 base model 開始,重新設計 architecture 與 training recipe。核心變化是首次在 GLM 系列採用 sparse attention 與 linear attention 組成的 hybrid architecture:linear attention 以 state modeling 捕捉 local dependencies,sparse attention 則透過 lightweight indexer 擷取 global context,藉此降低長上下文服務的計算成本,同時維持 long-context 能力。模型也導入 Manifold-Constrained Hyper-Connections(mHC),並使用最新的 30T-token multimodal pre-training corpus。 相較 GLM-5.3,Z.ai 文件列出的效率改善包括: Attention computation 降低 3.01×,其他文件約寫為 3.0×。 KV cache size 降低 4.44×,其他文件約寫為 4.4×。 在 1M-token context 下,IndexPool 會以 weighted pooling 將 4 個 indexer key vectors 壓縮成 1 個。 GLM-5.3-Flash 架構在長序列下顯著降低每層 KV-Cache 大小與 Attention 計算負擔,表現優於 GLM-5.3。 與 GLM-4.5 series 相比,總參數量為 320B 對 355B,啟用參數則由 32B 降至 18B,layers 也由 92 層降至 45 層。來源指出,GLM-5.3-Flash 的 attention compute 是比較模型中最低,但 KV cache 仍略大於 Kimi-K3 與 DeepSeek-V4-Flash,因此並非所有效率指標都已達到最佳。 可支援本地部署的 frameworks 包括 SGLang、vLLM、TokenSpeed 與 KTransformers。Z.ai 也表示,這是首個採用 sparse attention 與 linear attention hybrid architecture 的 open-source frontier model。 API 與多模態能力 API model code 為 glm-5.3-flash。除了文字,它原生支援圖片、影片與檔案;圖片可透過 messages[].content[] 的 type: imageurl content block 傳入,imageurl.url 可接受 image URL 或 Base64 Data URL,也能放入多個 image blocks。官方建議設定如下: `json { "temperature": 1, "top_p": 0.95, "reasoning_effort": "max", "thinking.type": "enabled", "thinking.clear_thinking": false, "stream": true, "tool_stream": true } ` 其中 thinking.type 只支援 enabled,thinking 不能停用;若使用 streaming requests,官方建議同時啟用 stream: true 與 tool_stream: true。相關介面與能力文件包括: Chat Completion API Thinking Mode Streaming Output Function Calling Context Caching Structured Output 完整文件索引:llms.txt 在 GLM Coding Plan 中,模型已 fully available,quota 為 GLM-5.3 的 3 倍;off-peak hours,包括週末全天,points-based quota 只消耗標準 points 的 50%。可參考 Personal Plan 與 Team Plan。 Coding 與 benchmark 表現 Z.ai 宣稱 GLM-5.3-Flash 在 coding 與 agentic benchmarks 多數勝過 GLM-5.2,且接近 Claude Opus 4.8。Artificial Analysis Intelligence Index v4.1.1 的得分為 57,每項 task 的 discounted 價格為 $0.045;來源據此估計,同等 intelligence 過去約需 10 倍成本。 具體結果包括: DeepSWE v1.1:63.4,相較 GLM-5.2 的 46.2。 AutomationBench:48.8,相較 GLM-5.2 的 26.2。 Z.ai Code Bench v1.0 在 Claude Code 2.1.207 上執行時,各 effort level 均勝過 GLM-5.2;max effort 為 29.0,Claude Opus 4.8 為 29.5。 Z.ai Code Bench 被 Z.ai 描述為衡量 real-world coding performance 的評測,官方稱 GLM-5.3-Flash 在每個 effort level 都優於 GLM-5.2,…

Episode metadata supplied by the publisher feed · Published Aug 26, 2026

Embed this episode

Ready to play

Z.ai 發布 GLM-5.3-Flash:以 18B 啟用參數支援原生多模態與 1M-token 長上下文

0:00 10:37

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of EasyVibeCoding Podcast?

This episode is 10 minutes long.

When was this EasyVibeCoding Podcast episode published?

This episode was published on August 26, 2026.

Can I download this EasyVibeCoding Podcast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!