PyTorch's Combined Effort in Large Model Optimization // Michael Gschwind // #274 episode artwork

EPISODE · Nov 26, 2024 · 57 MIN

PyTorch's Combined Effort in Large Model Optimization // Michael Gschwind // #274

from MLOps.community · host Demetrios

Dr. Michael Gschwind is a Director / Principal Engineer for PyTorch at Meta Platforms. At Meta, he led the rollout of GPU Inference for production services.// MLOps Podcast #274 with Michael Gschwind, Software Engineer, Software Executive at Meta Platforms.// AbstractExplore the role in boosting model performance, on-device AI processing, and collaborations with tech giants like ARM and Apple. Michael shares his journey from gaming console accelerators to AI, emphasizing the power of community and innovation in driving advancements.// BioDr. Michael Gschwind is a Director / Principal Engineer for PyTorch at Meta Platforms. At Meta, he led the rollout of GPU Inference for production services. He led the development of MultiRay and Textray, the first deployment of LLMs at a scale exceeding a trillion queries per day shortly after its rollout. He created the strategy and led the implementation of PyTorch donation optimization with Better Transformers and Accelerated Transformers, bringing Flash Attention, PT2 compilation, and ExecuTorch into the mainstream for LLMs and GenAI models. Most recently, he led the enablement of large language models on-device AI with mobile and edge devices.// MLOps Swag/Merchhttps://mlops-community.myshopify.com/// Related LinksWebsite: https://en.m.wikipedia.org/wiki/Michael_Gschwind --------------- ✌️Connect With Us ✌️ -------------Join our Slack community: https://go.mlops.community/slackFollow us on Twitter: @mlopscommunitySign up for the next meetup: https://go.mlops.community/registerCatch all episodes, blogs, newsletters, and more: https://mlops.community/Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/Connect with Michael on LinkedIn: https://www.linkedin.com/in/michael-gschwind-3704222/?utm_source=share&utm_campaign=share_via&utm_content=profile&utm_medium=ios_appTimestamps:[00:00] Michael's preferred coffee[00:21] Takeaways[01:59] Please like, share, leave a review, and subscribe to our MLOps channels![02:10] Gaming to AI Accelerators[11:34] Torch Chat goals[18:53] Pytorch benchmarking and competitiveness[21:28] Optimizing MLOps models[24:52] GPU optimization tips[29:36] Cloud vs On-device AI[38:22] Abstraction across devices [42:29] PyTorch developer experience[45:33] AI and MLOps-related antipatterns[48:33] When to optimize[53:26] Efficient edge AI models[56:57] Wrap up

Episode metadata supplied by the publisher feed · Published Nov 26, 2024

Embed this episode

NOW PLAYING

PyTorch's Combined Effort in Large Model Optimization // Michael Gschwind // #274

0:00 57:44

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of MLOps.community?

This episode is 57 minutes long.

When was this MLOps.community episode published?

This episode was published on November 26, 2024.

Can I download this MLOps.community episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!