vLLM episode artwork

EPISODE · May 4, 2025 · 13 MIN

vLLM

from Large Language Model (LLM) Talk · host AI-Talk

vLLM is a high-throughput serving system for large language models. It addresses inefficient KV cache memory management in existing systems caused by fragmentation and lack of sharing, which limits batch size. vLLM uses PagedAttention, inspired by OS paging, to manage KV cache in non-contiguous blocks. This minimizes memory waste and enables flexible sharing, allowing vLLM to batch significantly more requests. As a result, vLLM achieves 2-4x higher throughput compared to state-of-the-art systems like FasterTransformer and Orca.

Episode metadata supplied by the publisher feed · Published May 4, 2025

Embed this episode

NOW PLAYING

vLLM

0:00 13:06

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Large Language Model (LLM) Talk?

This episode is 13 minutes long.

When was this Large Language Model (LLM) Talk episode published?

This episode was published on May 4, 2025.

Can I download this Large Language Model (LLM) Talk episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!