LLM Inference Speed (Tech Deep Dive) episode artwork

EPISODE · Oct 6, 2023 · 39 MIN

LLM Inference Speed (Tech Deep Dive)

from Thinking Machines: AI & Philosophy · host Daniel Reid Cahn

In this tech talk, we dive deep into the technical specifics around LLM inference.The big question is: Why are LLMs slow? How can they be faster? And might slow inference affect UX in the next generation of AI-powered software?We jump into:Is fast model inference the real moat for LLM companies?What are the implications of slow model inference on the future of decentralized and edge model inference?As demand rises, what will the latency/throughput tradeoff look like?What innovations on the horizon might massively speed up model inference?

Episode metadata supplied by the publisher feed · Published Oct 6, 2023

Embed this episode

Ready to play

LLM Inference Speed (Tech Deep Dive)

0:00 39:36

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Thinking Machines: AI & Philosophy?

This episode is 39 minutes long.

When was this Thinking Machines: AI & Philosophy episode published?

This episode was published on October 6, 2023.

Can I download this Thinking Machines: AI & Philosophy episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!