EPISODE · Jan 20, 2025 · 7 MIN
大规模Transformer模型推理的效率优化
from AI Podcast · host weedge
本播客深入探讨了如何高效地部署大型Transformer模型进行生成式推理,特别是在延迟敏感和长序列长度的场景下。我们将讨论模型并行策略、内存优化和低级优化技术,这些技术共同实现了在延迟和模型FLOPS利用率方面的新的帕累托前沿。
Embed this episode
NOW PLAYING
大规模Transformer模型推理的效率优化
0:00
7:07
1×
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.
Frequently Asked Questions
How long is this episode of AI Podcast?
This episode is 7 minutes long.
When was this AI Podcast episode published?
This episode was published on January 20, 2025.
Can I download this AI Podcast episode?
Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!