#089 - Changing How LLMs Scale - The AI Token Breakthrough with Subquadratic CTO Alex Whedon episode artwork

EPISODE · Jul 23, 2026 · 1H 9M

#089 - Changing How LLMs Scale - The AI Token Breakthrough with Subquadratic CTO Alex Whedon

from GAEA Talks · host GAEA Talks

Graeme Scott sits down with Alex Whedon, Co-Founder and CTO of Subquadratic and former Head of Generative AI at Tribe AI, for one of the most first-principles conversations GAEA Talks has recorded.The entire AI industry, Alex argues, is downstream from a single algorithm - the transformer - and its two fundamental flaws are quietly capping what AI can do. He breaks down quadratic compute scaling in plain terms (10x the input, 100x the compute), the memory wall where context can cost more than the model itself, and how Subquadratic's linear-scaling architecture claims to cut compute by up to 1,000x at extreme context lengths without sacrificing quality. Along the way he makes a bracing case that electricity, water, minerals and capital - not clever engineering - are the real limits on AI's growth, that the industry is being far too stingy with tokens, and that efficiency isn't a nice-to-have but an inevitability.In this episode:Why the whole AI space is downstream from one algorithmQuadratic scaling explained simply - 10x the input, 100x the computeThe memory wall, where context can need more memory than the modelWhat "subquadratic" means, and why linear scaling is the unlockA claimed ~1,000x compute reduction at 12 million tokensWhy transformers are the worst fit for data-heavy enterprise workThe real constraints on AI: electricity, water, minerals and capitalWhy we're "too stingy with the tokens"The DeepSeek lesson the incumbents ignoredWhy first to market is rarely bestAbout Alex Whedon: Co-Founder and CTO of Subquadratic, which emerged from stealth in May 2026 with $29M in seed funding and SubQ - described as the first frontier model built on a fully subquadratic Sparse Attention architecture, with a research context window of up to 12 million tokens. Previously Head of Generative AI at Tribe AI, leading 40+ enterprise implementations for companies including Anthropic, New Relic and Mars, and earlier an engineer at Meta and Instagram.GAEA Talks is the enterprise AI podcast for leaders navigating the age of artificial intelligence. New conversations every week.

Episode metadata supplied by the publisher feed · Published Jul 23, 2026

Embed this episode

Ready to play

#089 - Changing How LLMs Scale - The AI Token Breakthrough with Subquadratic CTO Alex Whedon

0:00 1:09:36

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of GAEA Talks?

This episode is 1 hour and 9 minutes long.

When was this GAEA Talks episode published?

This episode was published on July 23, 2026.

Can I download this GAEA Talks episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!