EPISODE · Jun 22, 2026 · 8 MIN
Cheaper Tokens, Bigger Bills
from YPO Technology Network AI Brief
The strategic AI question is no longer "which model do we use." It's "where does the model run, and who pays for the tokens." This week the AI inference startup Baseten raised roughly $1.5 billion at up to a $13 billion valuation for the unglamorous business of running other companies' models. Meanwhile token prices are collapsing about 10x a year, and enterprise AI bills are going up anyway. In this episode, Stephen Forte unpacks the inference economy and what it means for your business: The inference gold rush — why investors value the company that runs models more than many that build them, and why inference is 80-90% of a model's lifetime cost. The land grab — Amazon selling its Trainium chips to challenge Nvidia, and Alphabet's $84.75 billion raise to fund AI capex. The pricing paradox — "LLMflation" makes tokens ~10x cheaper a year, yet the Jevons paradox and the new "thinking tax" of reasoning models send total bills higher. The counter-move — open-weight models running locally on your own hardware, Apple's new "zero token cost" Core AI, and how to think about cloud vs. local as a cost-structure decision. Two concrete moves for the quarter — build multi-model routing, and budget for usage growth, not the falling unit price. Sources: Baseten ~$1.5B round at up to $13B — TechCrunch Amazon to sell Trainium chips externally — TechCrunch Alphabet $84.75B equity offering for AI — Intellectia LLMflation and falling inference costs — a16z Jevons paradox and rising enterprise AI spend — GUUTs / FinOps Apple Core AI at WWDC26 — Let's Data Science The AI Brief from the YPO Technology Network is a daily executive briefing on the AI developments that matter to business leaders. Hosted by Stephen Forte.
Embed this episode
NOW PLAYING
Cheaper Tokens, Bigger Bills
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.