#601 The AI Bottleneck Is No Longer GPUs. It’s Energy and Memory | Eugene Cheah episode artwork

EPISODE · May 25, 2026 · 45 MIN

#601 The AI Bottleneck Is No Longer GPUs. It’s Energy and Memory | Eugene Cheah

from The CTO Show with Mehmet Gonullu · host Mehmet Gonullu

In this episode of The CTO Show with Mehmet, Mehmet sits down with Eugene Cheah, CEO of Featherless AI. The AI bottleneck is no longer just GPU access. Power, memory, inference cost, and model reliability are becoming the real constraints.Eugene reframes the AI infrastructure debate away from a simple race for bigger models and more chips. The conversation connects energy capacity, HBM shortages, open source model adoption, linear attention architectures, and the enterprise need for predictable AI systems. It also challenges the assumption that the best AI strategy is always to use the largest available model.If you are building, investing in, or operating AI infrastructure, this conversation gives a clearer view of where AI economics, hardware constraints, and production reliability are heading.About the GuestEugene Cheah is the CEO of Featherless AI, an AI startup making open source AI models accessible through a single platform.Featherless AI started from AI research and optimization work around RWKV architecture, with a focus on reducing inference cost and making AI models more accessible. Eugene’s work sits at the intersection of open source AI, model efficiency, GPU infrastructure, HBM constraints, and inference optimization.He is well positioned to frame this shift because Featherless AI works directly on the infrastructure layer between developers, open models, and production inference.LinkedIn: https://www.linkedin.com/in/eugene-cheah-a47791126/Website: https://featherless.aiKey TakeawaysAI infrastructure constraints are shifting from GPU access to power, memory, and inference efficiency.HBM scarcity becomes more serious as models and context windows continue to grow.Bigger models do not solve the enterprise problem of reliable execution.Open source models are becoming strong enough to replace many closed model use cases.Fine-tuned smaller models can outperform frontier models on narrow enterprise tasks.Nvidia’s moat weakens when developers can move workloads across more hardware choices.Linear attention architectures matter because quadratic memory scaling is economically unsustainable.Enterprises value model control when closed providers change, deprecate, or restrict models too often.What You Will LearnThe real infrastructure bottlenecks behind AI deployment beyond GPU availability.How HBM pressure affects model size, context length, and inference economics.Why energy capacity can delay AI infrastructure even when chips are already available.How open source models are changing enterprise AI adoption and deployment control.Why smaller fine-tuned models can beat larger models on specific production tasks.When linear attention architectures reduce memory demand compared with transformer attention.What hardware choice, model portability, and local inference mean for AI infrastructure strategy.Episode Highlights00:00 — AI infrastructure moves beyond the GPU race03:30 — Nvidia, AMD, and Huawei follow different hardware strategies07:30 — Power becomes the first AI infrastructure bottleneck08:30 — HBM pressure exposes the memory constraint12:00 — AI follows the same pluralism as databases15:00 — Developers start with big models, then specialize18:30 — Transformer memory scaling becomes an economic problem23:30 — Hardware choice starts weakening platform lock-in29:30 — Reliability matters more than raw intelligence36:00 — Open source gives enterprises model control41:30 — Small models can now build real applicationsResources MentionedFeatherless AI: https://featherless.aiRWKV architecture: AI architecture referenced by Eugene as part of Featherless AI’s research backgroundListen NowAvailable on all major podcast platforms and YouTube.Connect with the ShowFollow The CTO Show with Mehmet for more conversations at the intersection of technology, startups, and venture capital.

Episode metadata supplied by the publisher feed · Published May 25, 2026

Embed this episode

NOW PLAYING

#601 The AI Bottleneck Is No Longer GPUs. It’s Energy and Memory | Eugene Cheah

0:00 45:37

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of The CTO Show with Mehmet Gonullu?

This episode is 45 minutes long.

When was this The CTO Show with Mehmet Gonullu episode published?

This episode was published on May 25, 2026.

Can I download this The CTO Show with Mehmet Gonullu episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!