Why Cost Per Million Tokens Is A Useless KPI? episode artwork

EPISODE · Sep 14, 2026 · 38 MIN

Why Cost Per Million Tokens Is A Useless KPI?

from MLOps.community · host Demetrios

A year ago, Palo Alto Networks built dashboards to track AI spend. Today those dashboards are useless, and the team that built them thinks that's the whole story.Recorded at FinOps X in San Diego, this conversation brings together Abhinav Lad, who leads cloud and AI finance at Palo Alto Networks, and Kuntal Patel, who runs the cloud engineering function behind it. They explain what happened when agents entered the picture, and AI stopped behaving like a service anyone could forecast.The short version: consumption went from linear to exponential almost overnight. Agents are goal-oriented rather than task-oriented, so they plan, call tools, verify, fail, retry, and keep looping until they hit the outcome, and every iteration is billable. So how do you run finance on top of that? Abhinav and Kuntal walk through the metrics that replaced their old forecasts: adoption rate, cost per user, AI as a percentage of revenue - and the budget limits that let engineering leaders choose between the newest model and a longer runway. They get into the open question of whether a cheaper model saves money or just burns more tokens thinking. They explain why an AI gateway became the control plane for cost and security at the same time, why retry caps belong in the design phase instead of the postmortem, and how FinOps starts to resemble product QA once the bill becomes the clearest signal that something is broken.They close on a warning worth sitting with: cost per million tokens is a number that means almost nothing on its own, and a value story built on it will point you somewhere you don't want to go.Palo Alto Networks: https://www.paloaltonetworks.comAbhinav Lad: https://www.linkedin.com/in/abhinav-ladKuntal Patel: https://www.linkedin.com/in/kuntalpatel35Alex Salkever: https://www.linkedin.com/in/alexsalkeverTimestamps:[0:00] Intro[1:00] Who runs FinOps for AI at Palo Alto Networks[2:10] Last year's AI dashboards are already useless[4:26] Agents turned linear forecasts exponential[7:27] Three traits that make agents expensive[8:34] The hidden bill: RAG, vectors and egress[9:16] Cost per user and adoption rate[11:21] Giving engineering leaders a budget[12:07] Using DORA metrics to prove value[13:53] Where DORA stops fitting AI[16:20] Does the cheaper model actually save money[17:57] Why you need an AI gateway[20:05] Inside Prisma AIRS[21:00] Three cost models for three use cases[22:52] Forecasting lessons from Electronic Arts[24:03] Runaway agents and endless loops[25:59] Capping retries before they burn cash[28:06] Writing cost policy at design time[29:01] When FinOps becomes product QA[32:17] Explaining AI spend to the C-suite[34:51] Valuing AI beyond engineering[37:04] Crawl, walk, run: where they are today[38:20] Why cost per million tokens is meaningless[39:26] Closing thoughts

Episode metadata supplied by the publisher feed · Published Sep 14, 2026

Embed this episode

Ready to play

Why Cost Per Million Tokens Is A Useless KPI?

0:00 38:49

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of MLOps.community?

This episode is 38 minutes long.

When was this MLOps.community episode published?

This episode was published on September 14, 2026.

Can I download this MLOps.community episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!