Why Your AI Bill Will Double Before It Gets Better episode artwork

EPISODE · Aug 3, 2026 · 32 MIN

Why Your AI Bill Will Double Before It Gets Better

from MLOps.community · host Demetrios

In this episode, we're joined by Josh Collier, FinOps Lead at Superhuman (formerly Grammarly), to explore what it really costs to run AI at scale and why the rules of the game changed faster than anyone expected.We discuss how AI token costs dropped 80% in two years, why that trend has sharply reversed with frontier models doubling in price, and how Josh rebuilt a single LLM workflow that cost $400k a month down to $80k by rethinking the architecture. He also shares how a cost calculator built in 15 minutes transformed the way his team estimates spend before running experiments, and why research-led optimization is the only kind that works without degrading the product.Along the way, we cover hidden costs most teams miss, the trade-off between Azure reserved capacity and OpenAI Priority Processing, why fixed subscription pricing is broken in an AI-native world, vendor lock-in risk, and what OpenAI's Guaranteed Capacity announcement really signals about where vendor relationships are heading next.Superhuman: https://superhuman.comJosh Collier: https://www.linkedin.com/in/josh-collier-945b7029/Demetrios: https://www.linkedin.com/in/dpbrinkmTimestamps:[00:00] OpenAI Guaranteed Capacity: what's really going on[01:04] Josh's path into AI FinOps[02:48] Token costs: the 80% price drop[04:16] Why costs will only go up[05:06] External LLMs as financial risk[07:16] Why subscription pricing is dead[08:22] The data residency fee nobody notices[09:33] The cost calculator built in 15 minutes[10:24] How it changed dev team speed[13:00] Tracking costs by service and team[15:33] $400k workflow rebuilt for $80k[17:13] Why only research can optimize tokens[20:00] Speculative decoding win[23:11] One bad query, $40k gone[26:00] Why Azure PTU was exhausting[28:59] Shadow traffic load testing[29:07] Priority processing: no brainer[31:10] Guaranteed capacity: lock-in signal?[32:18] The danger of multi-year AI deals[33:28] Vendor-agnostic proxy as exit strategy

Episode metadata supplied by the publisher feed · Published Aug 3, 2026

Embed this episode

Ready to play

Why Your AI Bill Will Double Before It Gets Better

0:00 32:24

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of MLOps.community?

This episode is 32 minutes long.

When was this MLOps.community episode published?

This episode was published on August 3, 2026.

Can I download this MLOps.community episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!