大规模Transformer模型推理的效率优化 episode artwork

EPISODE · Jan 20, 2025 · 7 MIN

大规模Transformer模型推理的效率优化

from AI Podcast · host weedge

本播客深入探讨了如何高效地部署大型Transformer模型进行生成式推理,特别是在延迟敏感和长序列长度的场景下。我们将讨论模型并行策略、内存优化和低级优化技术,这些技术共同实现了在延迟和模型FLOPS利用率方面的新的帕累托前沿。

Episode metadata supplied by the publisher feed · Published Jan 20, 2025

Embed this episode

NOW PLAYING

大规模Transformer模型推理的效率优化

0:00 7:07

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of AI Podcast?

This episode is 7 minutes long.

When was this AI Podcast episode published?

This episode was published on January 20, 2025.

Can I download this AI Podcast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!