PODCAST · technology
Humans of Reliability
by Rootly
Behind every reliable software system, there are people working hard to keep it online. Humans of Reliability is a series that spotlights the engineers, leaders, and innovators at the heart of incident management and system reliability. Through candid conversations, we explore the challenges, lessons, and personal journeys of those navigating complex technical landscapes to ensure the systems we rely on run smoothly. From unforgettable incident stories to favorite tools, workflows, and hobbies, Humans of Reliability uncovers the human side of technology—offering insights and inspiration for anyone passionate about building and maintaining resilient systems.https://rootly.com/humans-of-reliability
-
33
Your AI agents are lost: give them a graph w/ Anthony Alcaraz (AWS)
The race to build capable AI agents often focuses on model intelligence. Anthony Alcaraz argues that the harder, more durable problem is context: giving agents the right information, relationships, tools, and constraints at the moment they need them.Anthony is a Senior AI/ML Portfolio Growth Manager at AWS and co-author of O'Reilly's Agentic GraphRAG. He explains how graph-based architectures can improve retrieval, memory, planning, and multi-step reasoning while capturing the tribal knowledge that conventional enterprise systems leave behind.The conversation also looks beyond architecture. Anthony describes self-improving agent systems, feedback loops that connect technical performance to business outcomes, and a future where small teams manage large fleets of agents without losing sight of customers or value.
-
32
AI vs. AI: From Alert Fatigue to Agentic Cybersecurity
AI is giving attackers the ability to scale phishing, discover vulnerabilities, and build increasingly autonomous malware. But it is also helping defenders investigate alerts, eliminate noise, and respond before threats become incidents.Nir Soudry, Head of R&D at 7AI, joins Humans of Reliability to explore the rise of agentic cybersecurity. He explains why humans must remain accountable for AI decisions, how security teams can introduce agents gradually, and why adopting AI should be treated as an ongoing partnership rather than a simple software installation.We also discuss AI-generated code and its impact on reliability, 7AI’s claim that automated investigations can remove roughly 90% of alert noise, and the industry’s evolution from reactive security toward proactive, policy-driven prevention.Will generative AI create more incidents—or help us stop them before they happen?
-
31
Stop shipping features for a quarter, it might save you w/ Monday.com co-founder
When incidents pile up fast enough, every part of the company bleeds: support is fielding angry customers, AEs are on apology calls, and engineering is burning cycles on retrospectives instead of shipping. For Eran Kampf (VP of Engineering at Twingate) where the product is the network, that was the moment he made a call most engineering leaders won't: stop all feature work for a quarter and fix reliability.In this episode, Eran walks through what lead to that decision, why best-in-class engineering practices (full test coverage, Terraform provisioning, comprehensive monitoring) turned out to be just the baseline for four nines and what it took to rebuild trust internally and with customers.
-
30
"Hey Claude, where's my database?": How an AI agent nuked production w/ Alexey Grigorev
"Hey Claude, where's my database?" The answer came back: "Oh, sorry, your database is gone." Alexey Grigorev, founder of DataTalks.Club and creator of the Zoomcamp courses that have taught data and ML engineering to over 100,000 people, knows Terraform well enough to cover it in his curriculum. That didn't stop a chain of small, reasonable-sounding decisions from ending in an AI agent running terraform destroy against his live production database.
-
29
Every pilot is ready for engine failure: are your engineers? w/ Hamed Silatani (Uptime Labs)
Every pilot who's never had an engine failure is still ready for one. The same can't be said for most software engineers facing their first major incident. Hamed Silatani, co-founder and CEO of Uptime Labs, and former Head of Reliability Engineering at IG Group, has spent two decades watching engineers learn incident response the hard way: alone, under pressure, with no training. A moment in a hospital delivery room, where a team of nurses responded to a crisis with clockwork precision, crystallized what he'd long suspected: IT spends heavily on tooling and process but almost nothing on training the humans who have to use them under fire. In this conversation, Hamed breaks down the three skill categories that matter in incident response, explains why communication is always the first pain point leaders name, and makes the case that AI-driven complexity is about to make human incident skills more important, not less.
-
28
LLM Observability: Lessons From MLOps w/ Maria Vechtomova (Cauchy)
For nine years, Maria Vechtomova was shouting about monitoring. Nobody cared, until LLMs arrived. As co-founder of Cauchy, Databricks MVP, and one of the most followed voices in MLOps, Maria has watched the field evolve from hand-built experiment trackers to today's flood of observability tools, and her central claim might surprise you: globally, nothing has changed. The fundamentals are the same: track your code, data, and models so you can roll back when something breaks. What did change is the surface area. Tools, prompts, embeddings, agents, every component shifts behavior unpredictably, and business metrics often become the only signal left. Maria gets into why most teams still can't roll back cleanly, why 40% of ML projects ship with no monitoring at all, and why she believes the next era of MLOps will be its biggest yet.
-
27
The Golden Hour: Why the First 15 Minutes of an Incident Decide Everything w/ Gandhi M. N. Kumar (Twillio)
Most incident response advice focuses on tools, alerts, and post-mortems. Gandhi Mathi Nathan Kumar, Principal Incident Commander at Twilio, with 14 years running calls that have pulled in up to 100 responders, argues the work that actually matters happens in the first 15 minutes. In this episode, Gandhi walks through what he calls the golden hour: the window where you decide whether you know what's broken, who belongs on the call, and whether to chase the root cause or reach for redundancy. He gets into why mitigation has to come before diagnosis, why customers trust your status page more than your engineers, and why he once sat with a stopwatch counting how many clicks it took to declare an incident. Along the way: the human side leaders keep underinvesting in, the math of on-call fatigue, and where AI is actually pulling weight in the incident commander seat.
We're indexing this podcast's transcripts for the first time — this can take a minute or two. We'll show results as soon as they're ready.
No matches for "" in this podcast's transcripts.
No topics indexed yet for this podcast.
Loading reviews...
ABOUT THIS SHOW
Behind every reliable software system, there are people working hard to keep it online. Humans of Reliability is a series that spotlights the engineers, leaders, and innovators at the heart of incident management and system reliability. Through candid conversations, we explore the challenges, lessons, and personal journeys of those navigating complex technical landscapes to ensure the systems we rely on run smoothly. From unforgettable incident stories to favorite tools, workflows, and hobbies, Humans of Reliability uncovers the human side of technology—offering insights and inspiration for anyone passionate about building and maintaining resilient systems.https://rootly.com/humans-of-reliability
HOSTED BY
Rootly
CATEGORIES
Loading similar podcasts...