Compute scarcity is an engineering problem episode artwork

EPISODE · Jun 30, 2026 · 7 MIN

Compute scarcity is an engineering problem

from Air Street Press

Angelos Perivolaropoulos, a research engineer at ElevenLabs, on turning GPU scarcity into an inference-engineering problem: how to serve far more users on the same hardware, from batching to frontier architecture changes. Recorded at RAAIS 2026.00:00 Introduction: ElevenLabs and the GPU squeeze00:38 The question: how to scale when you can't add capacity01:11 About Angelos: Scribe, speech-to-text and text-to-speech01:56 GPU scarcity meets exponential demand02:44 What a token actually costs: compute vs memory bandwidth03:38 Prefill, decode and the KV cache05:53 Batching and continuous batching (1 → 15 users/GPU)08:37 FP8 quantization and quantize-aware training (→ 20)11:29 Speculative decoding and multi-token prediction (→ 28)15:13 Compressing the KV cache: TurboQuant and distillation (→ 70)17:27 Frontier architectures: MLA, linear attention, state-space (→ 140)20:39 Trade-offs: nothing is free22:03 Q&A: papers vs production, token subsidies, TTS evals

Episode metadata supplied by the publisher feed · Published Jun 30, 2026

Embed this episode

NOW PLAYING

Compute scarcity is an engineering problem

0:00 7:44

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Air Street Press?

This episode is 7 minutes long.

When was this Air Street Press episode published?

This episode was published on June 30, 2026.

Can I download this Air Street Press episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!