MiniMax 推出開放權重 MiniMax-Music3,生成最長五分鐘完整歌曲 episode artwork

EPISODE · Aug 13, 2026 · 2 MIN

MiniMax 推出開放權重 MiniMax-Music3,生成最長五分鐘完整歌曲

from EasyVibeCoding Podcast

MiniMax 推出開放權重 MiniMax-Music3,生成最長五分鐘完整歌曲。 MiniMax Music 3 的發表橫幅,背景為帶有羽毛紋理的紫色漸層,中央以白色字樣標示主標題與「Next-Generation Open-Weights Production-Ready & Versatile Music Model」副標題。 模型定位 MiniMax 透過 Hugging Face、GitHub 與 魔搭 分享模型。它可根據歌詞與音樂描述,生成最長五分鐘、32 kHz、16-bit 立體聲 WAV 完整歌曲,維持主題、節奏、人聲識別與編曲發展。 紫色背景宣傳海報,中央醒目顯示白色大字「MiniMax Music 3」,下方標註「Next-Generation Open-Weights Production-Ready & Versatile Music Model」產品定位說明。 控制方式 歌詞可加入 [Intro]、[Verse]、[Chorus]、[Bridge] 等段落標籤;音樂描述則指定曲風、情緒、人聲、樂器與製作風格。官方建議使用包含 Global Metadata、Vocal Details、Arrangement 的 Structured Caption,細化歌曲各段落的變化。 混合語言模型架構圖,上方顯示 Flow-Matching 與 Flow VAE Decoder 的 Synthesis 區塊,中間包含 Global LLM、Local LLM 與 Hidden state 和 Stop Token,下方則為 Structured Caption 與 Lyrics 的 Input Conditions 區塊。 技術架構 模型以 8B Global LLM 處理長距離結構,0.6B Local LLM 還原每個 frame 的聲學細節,再融合 hidden states,經過 Flow Matching(2.4B)與 Flow-VAE Decoder(123M)合成音訊。訓練 tokenizer 使用八層 RVQ,第一層有 16,384 個語意 codebook entries,其餘七層各有 1,024 個聲學 entries。 部署與操作 官方文件提供以下流程: 下載模型: `bash hf download MiniMaxAI/MiniMax-Music3 --local-dir /path/to/minimax_ttm ` 啟動 SGLang-Omni: `bash sgl-omni serve --model-path MiniMaxAI/MiniMax-Music3 --port 8000 ` 呼叫 API 生成音訊: `bash curl http://127.0.0.1:8000/v1/audio/speech \ -H 'Content-Type: application/json' \ -d '{ "model": "MiniMaxAI/MiniMax-Music3", "input": "[Verse]\nMorning light filtering through the pine\n[Chorus]\nSoftly the world begins to breathe", "instructions": "A warm acoustic pop song with intimate female vocals, fingerpicked guitar, soft piano, and a gradual emotional build into a wide final chorus.", "response_format": "wav", "seed": 7, "maxnewtokens": 750, "stream": false }' \ --output minimax_music3.wav ` Prompt 強化與限制 music-caption-rewriter skill 可將簡短描述擴展成 Structured Caption,安裝指令為: `bash npx skills add MiniMax-AI/MiniMax-Music3 --skill music-caption-rewriter ` 目前推論需要兩張 CUDA GPU,僅支援非串流生成;文字 prompt 上限為 5,000 tokens、音訊上限為 9,000 acoustic frames,段落標籤與描述只提供生成控制,無法保證完全符合指定速度、調性、歌詞或結構。原文:https://easyvibecoding.app/curated/2958-minimax-launches-open-weight-music-model-five-minute-songs

Episode metadata supplied by the publisher feed · Published Aug 13, 2026

Embed this episode

Ready to play

MiniMax 推出開放權重 MiniMax-Music3,生成最長五分鐘完整歌曲

0:00 2:59

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of EasyVibeCoding Podcast?

This episode is 2 minutes long.

When was this EasyVibeCoding Podcast episode published?

This episode was published on August 13, 2026.

Can I download this EasyVibeCoding Podcast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!