LLM 推理的效率前緣:延遲與成本的殘酷取捨 The efficient frontier of LLM inference episode artwork

EPISODE · Sep 3, 2026 · 6 MIN

LLM 推理的效率前緣:延遲與成本的殘酷取捨 The efficient frontier of LLM inference

from 蝦生實驗室

📝 本集重點:• 效率前緣原本是 Markowitz 投資組合理論的概念,風險跟報酬之間存在一條最優曲線,你只能沿著曲線移動,跳不出去。套在推理上就是:同一套硬體、同一個模型,你把單一使用者的延遲壓得越低,每張 GPU 能榨出的總 token 數就越少,單位...• 對,這個區分超重要。speculative decoding、quantization、更好的 kernel、編譯層的優化,這些才是真的把前緣往外推。至於「我們家 API 超快」——等一下,你先講清楚你 batch 開多少,不然那叫行銷,不...• 而且這不是工程師偷懶,是物理限制。本質上就是 batching。GPU 是吞吐量怪獸,它最怕的就是你只餵它一個請求。跑 batch size 1 的時候,大部分時間都耗在把模型權重從記憶體搬進運算單元,算力根本在發呆,利用率低到可笑。🤓• 而且不同應用要的位置完全不一樣。語音助理那種,使用者在等聲音出來,首 token 延遲多半秒體驗就毀了,你必須站在低延遲那端,貴也得吞。可是批次資料處理、半夜排程做摘要,慢幾秒根本沒人在乎,就該站在高吞吐那端把成本壓到最低。

Episode metadata supplied by the publisher feed · Published Sep 3, 2026

Embed this episode

Ready to play

LLM 推理的效率前緣:延遲與成本的殘酷取捨 The efficient frontier of LLM inference

0:00 6:54

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of 蝦生實驗室?

This episode is 6 minutes long.

When was this 蝦生實驗室 episode published?

This episode was published on September 3, 2026.

Can I download this 蝦生實驗室 episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!