Shopify 以 Qwen3.5-0.8B 將 Buyer Profile 產能提高 36 倍;Flow 另用 Qwen3-32B 與 Tangle 建立 production 回饋迴圈 episode artwork

EPISODE · Sep 2, 2026 · 2 MIN

Shopify 以 Qwen3.5-0.8B 將 Buyer Profile 產能提高 36 倍;Flow 另用 Qwen3-32B 與 Tangle 建立 production 回饋迴圈

from EasyVibeCoding Podcast · host tobi lutke

Shopify 以 Qwen3.5-0.8B 將 Buyer Profile 產能提高 36 倍;Flow 另用 Qwen3-32B 與 Tangle 建立 production 回饋迴圈。Buyer Profile 結果 Tobi Lutke 分享的 Buyer Profile slide 顯示,Qwen3.5-0.8B student 在特定任務拿到 84.6,較 GPT-5.6-sol xhigh teacher 的 83.0 高 1.6 個百分點;原本 9.1K-token 的 system prompt 壓到 1.1K tokens,約剩八分之一。在 100 張 H100 上,處理量從每日 2M profiles 成長為 72M,成長為 36 倍。過程中,模型分數由 7 月 23 日的 75.3(29K samples),提升至 7 月 27 日的 78.1(42K samples),再到 7 月 30 日的 84.6(54K samples);Qwen3.5-2B production baseline 則為 77.2。Qwen3.5-0.8B 在 54k 樣本微調後 Judge score 達到 84.6,超越 GPT-5.6-sol (xhigh) Teacher 的 83.0 與前代生產模型 Qwen3.5-2B 的 77.2,同時 system prompt 壓縮至 1.1K tokens,每日產出 Profile 提升至 72M。Flow production loop Shopify 另一篇 Flow case study 使用的是 Qwen3-32B,不是 Buyer Profile 的 Qwen3.5-0.8B;兩案使用的資料、評估器、流量與指標都不同,不能把分數或效能數字互相比高低。Flow 將商家的自然語言轉成會呼叫工具的自動化流程。團隊以 300 個人工設計樣本評估語意正確性、語法與延遲;把目標格式從巢狀 JSON DSL 改為 Python 後,語法正確性提高 22 個百分點、語意正確性提高 13 個百分點。模型訓練需在兩個 H200 節點上執行 12 小時。Qwen3-32B fine-tuned 在 LLM Judge Overall Pass Rate 指標上自 Jan 的 22.9% 升至 April 的 56.1%,並在 April 領先 Closed-Source Frontier Model 的 51.0%。真實流量驗證 Flow 首次只開放 1% 流量時,即使離線基準測試與原方案相當,實際工作流啟用率仍低 35%,顯示離線測試可能漏掉生產環境的特定失敗情境。Shopify 以數百筆人工標註對話校準 LLM 評分器,接著每週匯入 production 對話並評分,把高品質樣本送進訓練、隔離低品質樣本、找出缺口、重新訓練再部署。Shopify 表示這套 flywheel 在兩週內補上 offline 與 production 的品質落差;目前上線的 32B Flow agent 相較它取代的 frontier model 快 2.2 倍、成本低 68%。官方沒有公開兩週的起訖日期,也沒有提供成本前後值或計算式,因此這些仍是第一方 production 結果。Tangle 基礎設施 Tangle 是 Shopify ML 與軟體工程師打造的平台無關、可重現的 AI pipeline 系統,不是 learning algorithm。其視覺編輯器把可重用元件組成 pipeline,元件可在容器內使用 Python、Java、shell、Ruby、C++ 或 JavaScript/TypeScript;content-based caching 會重用未受變更影響的 upstream outputs,讓原本約 10 小時的流程在只改動單一元件時縮短至約 20 分鐘。公開的 TangleML/tangle backend 程式庫採 Apache License 2.0。Tangle Pipeline 的運作流程圖,大框內包含 Data Pipeline 連至 Training(2x H200 / 12 hours)、Evaluation LLM Judge 與 Deploy CentML,右側外接 HuggingFace Datasets & Models 以及下方 CometML Experiment Tracking。重現邊界 Tangle 提供的是資料、訓練與評估的迭代編排;若沒有評估器、資料分流政策、隔離流程、重新訓練與部署門檻,模型不會自行改善。Buyer Profile 尚未提供評分提示詞、測試切分、信賴區間、訓練成本、teacher 輸出、資料集、可重現的公開 checkpoint 或完整做法,因此 84.6 分仍是第一方報告,無法只靠目前公開的資料獨立重現。原文:https://easyvibecoding.app/curated/3237-shopify-0-8b-buyer-profile-tangle-production-feedback-loop

Episode metadata supplied by the publisher feed · Published Sep 2, 2026

Embed this episode

Ready to play

Shopify 以 Qwen3.5-0.8B 將 Buyer Profile 產能提高 36 倍;Flow 另用 Qwen3-32B 與 Tangle 建立 production 回饋迴圈

0:00 2:07

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of EasyVibeCoding Podcast?

This episode is 2 minutes long.

When was this EasyVibeCoding Podcast episode published?

This episode was published on September 2, 2026.

Can I download this EasyVibeCoding Podcast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!