EPISODE · Jun 13, 2026 · 6 MIN
How One CTO Used Observability to Find a Memory Leak Saving Millions
from Tech Leadership with Fexingo: Engineering Managers, CTOs, and Technical Leadership Conversations · host Fexingo
Lucas and Luna dive into a real-world story from a mid-sized SaaS company where the CTO used observability data — not intuition — to locate a subtle memory leak that was silently inflating their cloud bill by $40,000 a month. They walk through the specific signals: a gradual tail-latency creep, a steady rise in container restarts, and a single anomalous metric in their Prometheus dashboard that everyone had dismissed as noise. The hosts explain how the team built a custom alert based on RSS (residual sum of squares) to detect the leak before it impacted customers, and why this case proves that observability isn't just for reliability — it's a cost-control lever. They also discuss the cultural shift required to get engineers to treat metrics as money, and how the CTO tied observability to the P&L without creating a blame culture. #Observability #MemoryLeak #CloudCost #SaaS #Prometheus #CTO #EngineeringLeadership #SiteReliability #SRE #Monitoring #Alerting #CostOptimization #TailLatency #ContainerRestarts #ResidualSumOfSquares #FexingoBusiness #BusinessPodcast #Technology Keep every episode free: buymeacoffee.com/fexingo
Embed this episode
NOW PLAYING
How One CTO Used Observability to Find a Memory Leak Saving Millions
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.