Unsloth Desktop 將訓練與部署留在本機,VRAM 減少 70% episode artwork

EPISODE · Aug 12, 2026 · 8 MIN

Unsloth Desktop 將訓練與部署留在本機,VRAM 減少 70%

from EasyVibeCoding Podcast · host Unsloth AI

Unsloth Desktop 將訓練與部署留在本機,VRAM 減少 70%。 分享重點 UnslothAI 於 2026 年 8 月 11 日介紹 Unsloth Desktop,定位為免費、open-source 的跨平台桌面應用程式。它可在 macOS、Windows、Linux 以及多種 CPU、GPU 硬體上運作,讓使用者不必把模型、資料與推論流程交給雲端,就能在自己的電腦上聊天、訓練、執行程式碼與部署服務。官方入口為 unsloth.ai、GitHub、Unsloth Desktop 文件,完整文件索引見 llms.txt 與 Markdown 文件。 這項分享的核心價值,是把原本分散在模型下載、量化、推論、fine-tuning、Agent 工具串接與遠端部署的工作,集中到一個可在本機執行的介面中。官方宣稱訓練速度可達 2 倍、VRAM 使用量減少 70%,tool calls 的準確率最多提升 50%;這些數字均為 Unsloth 的產品宣稱,實際結果仍會受到 model、硬體、量化設定與工作負載影響。 跨平台與模型支援 Unsloth Desktop Beta 支援 macOS、Windows、Linux,並涵蓋 LLM、diffusion image/video、MLX、GGUF、embedding 與 audio models。下載入口分為 macOS、Windows 及 Linux and WSL;安裝文件分別是 https://unsloth.ai/docs/get-started/install/mac.md、https://unsloth.ai/docs/get-started/install/windows-installation.md 與 https://unsloth.ai/docs/get-started/install/linux.md。 首次使用時,使用者可從「Select model」選單或「Model hub」挑選符合裝置能力的 model 與 quantization,下載完成後即可開始聊天,不需要額外設定。官方列出的模型包括 Kimi K3、MiniMax-H3、Qwen3.8、Muse Glimmer、DeepSeek-V4、Gemma 4、Qwen-AgentWorld、Ornith、Kimi K2.7 Code、MiniMax M3、GLM-5.2、DiffusionGemma、Qwen3.6 與 Gemma 4;其中 GLM-5.2 被描述為 Z.ai 的 744B-parameter、1M-context open model,可透過 Dynamic GGUFs 在本機執行。 Unsloth 也宣稱支援 CPU、NVIDIA、AMD、Intel 與 Mac GPU,以及 multi-GPU 配置。AMD GPU 可在 Windows、WSL 與 Linux 上進行 training、RL、chat 與 deployment;GGUF 則支援 GPU layer placement、MoE experts offload、multi-GPU 與 Tensor Parallelism,相關變更見 PR #6414。Vulkan 目前只用於加速 GGUF inference,不會提供 training;training 仍需要受支援的 PyTorch 或 MLX backend。較舊硬體可能支援不佳,使用前應查閱 compatible GPUs 與 AMD guide。 桌面端 AI 執行與本地模型管理工具 Unsloth Desktop 介面功能展示 本機聊天與 Agent 整合 Unsloth Desktop 能把 local LLM 接到 Claude Code、Codex、OpenCode、OpenClaw、Hermes 等 agent 工具。這項功能透過 model swapping 運作:Claude Code 或 Codex 可以保留目前的 model,也可以把 Unsloth model 當作 local subagent。相關文件包括 Claude Code、Codex、MCP 與 Unsloth Start。 啟動 Unsloth、載入 model、開啟 project folder 後,可以依序執行: `bash unsloth start claude unsloth start codex unsloth start hermes unsloth start openclaw unsloth start opencode unsloth start claude --as-subagent --model unsloth/model-GGUF:quant ` unsloth start 也能透過 OpenAI-compatible 或 Anthropic-compatible API 連接 agents;MCP endpoint 可管理 models、training、recipes、checkpoints 與 exports,相關變更見 PR #7191。文件列出的相關工具頁面包括 Claude Code、Codex、Hermes Agent、OpenClaw、Unsloth API 與 OpenCode,其素材連結則分別是 Claude Code 素材、Codex 素材、Hermes Agent 素材、OpenClaw 素材、Unsloth API 素材 與 OpenCode 素材。 圓形的綠色貼紙圖示,中央帶有白底黑線的樹懶圖案,右下角有一處微微掀起的銀色邊角。 工具呼叫與程式碼執行 Unsloth Desktop 提供 self-healing tool calls,會偵測失敗、嘗試修復並重新執行;Bash 與 Python 則在 secure sandbox 中執行,讓 model 能執行程式碼、檢查結果並完成實際任務。官方稱 tool-calling accuracy 最多提升 50%,但同時承認 web search、code execution 與 tool-call healing 會增加推論時間;關閉這些功能後,速度應接近其他 llama.cpp app,若仍然緩慢,官方建議提交 GitHub issue。 它也整合 private、unlimited web search、deep research、RAG、MCP、image/video generation 與 TTS。音訊能力可完全在本機進行 generate、fine-tune 或 transcribe,涵蓋 text-to-speech、speech-to-text、Whisper 與 Qwen3-ASR;影像與影片則支援 MiniMax-H3、FLUX、Z-Image、Wan、LTX 及 fine-tuned LoRA adapters,並能進行 transform、inpaint、extend、upscale、reference 與既有影像編輯。相關文件為 audio fine-tuning、image fine-tuning 與 advanced web search。 Unsloth 的介面視窗,左側欄顯示導覽選單與近期對話紀錄,中央顯示 Muse-Glimmer-30B-GGUF 模型與 token 使用量,下方對話輸入框已展開功能選單,包含 Web search、Code、Deep research 與 MCP 等選項。 訓練與匯出 Unsloth Studio 的 web UI 與 Unsloth Core code-based 介面,和 Desktop 共同構成三種使用方式。使用者可以直接提供 PDF、CSV、JSON 或 DOCX 建立 datasets,執行 LoRA、QLoRA、full fine tuning、pretraining、reinforcement learning、GRPO、DPO 與 FP8 訓練,也能為 SDXL、FLUX.2、Qwen-Image、Z-Image 訓練 LoRA adapters。 官方列出的效能宣稱包括: 一般 fine-tuning 速度成長為兩倍,VRAM 使用量減少 70%。 MoE LLM training 速度成長為 12 倍,VRAM 使用量減少 35%,涵蓋 DeepSeek、GLM、Qwen 與 gpt-oss。 Embedding fine-tuning 速度成長為 1.8–3.3 倍。 新 batching algorithms 讓 context RL 可使用比其他 setups 長 7 倍的 context。 新 RoPE、MLP Triton Kernels,以及 Padding Free + Packing,宣稱 training 速度成長為 3 倍、VRAM 使用…

Episode metadata supplied by the publisher feed · Published Aug 12, 2026

Embed this episode

Ready to play

Unsloth Desktop 將訓練與部署留在本機,VRAM 減少 70%

0:00 8:59

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of EasyVibeCoding Podcast?

This episode is 8 minutes long.

When was this EasyVibeCoding Podcast episode published?

This episode was published on August 12, 2026.

Can I download this EasyVibeCoding Podcast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!