Why AI Companies Keep Wasting Billions On Idle GPUs episode artwork

EPISODE · Jul 14, 2026 · 41 MIN

Why AI Companies Keep Wasting Billions On Idle GPUs

from Thinking On Paper · host Mark Fielding and Jeremy Gilbertson

AI companies are splashing billions on more GPUs before they use the ones they already have. Many sit idle, others use a fraction of their potential. For the hyperscalers, cloud providers and AI companies, buying more GPUs does not solve much if the GPUs already in the rack spend too much time waiting for the rest of the system.The next battle in the AI infrastructure wars is efficiency: GPU utilization, NICs, memory movement, networking, CPUs and software that keeps expensive hardware doing useful work.Moshe Tanach runs NeuReality AI, a company building infrastructure for AI inference. In this conversation, he explains why the next wave of AI infrastructure depends on GPU utilization, NICs, memory movement, networking and software efficiency, not just buying more chips.This is a technical conversation about AI inference, infrastructure efficiency and the hardware layer beneath ChatGPT, Claude, Gemini and the AI tools now being built into business software.Please enjoy the show.Thinking on Paper is a technology podcast about AI, Space, quantum computing, science, and the systems shaping your life. 🏠 ⁠Buy us a beer on Substack⁠🫵 C⁠hoose your own technology adventure ⁠📺  ⁠Watch our beautiful faces on YouTube ⁠🎧 R⁠emember Steve Jobs on APPLE⁠📺 ⁠Get clips and exclusive videos on Instagram ⁠--Chapters(00:00) The GPU challenge(01:27) Training Vs Inference(05:45) Memory & Keeping The GPU Busy(07:31) Deep Seek, Blackwell & Ruben(09:03) How Ripe Is CPU For Reinvention?(11:03) AI Agents Run On CPUs(12:51) Do We Need All The Computation?(14:35) GPU Utilization Rates(16:50) Tokens In A Context Window(20:59) Data, Knowledge & Wisdom(23:56) Why Are Your GPUs Busy? (A Truck Analogy)(29:07) Hyperscalers(31:59) The Decode Phase(36:47) What More Efficient GPUs Means For The User

Episode metadata supplied by the publisher feed · Published Jul 14, 2026

Embed this episode

NOW PLAYING

Why AI Companies Keep Wasting Billions On Idle GPUs

0:00 41:19

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Thinking On Paper?

This episode is 41 minutes long.

When was this Thinking On Paper episode published?

This episode was published on July 14, 2026.

Can I download this Thinking On Paper episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!