EPISODE · Aug 5, 2026 · 30 MIN
Building AI Software That You Can Actually Benchmark | David Haddad
from Advocate Insurance Desk · host Advocate Technologies
Every vendor in commercial insurance now says they use AI. Almost none of them will tell a buyer which part of an answer was calculated and which part was generated. David Haddad, head of product engineering at Advocate Technologies and the builder of the World Insurance Model covered in Episode 21, joins Katie Dowson to explain why that distinction is the whole thing, and why the model itself is the least interesting part of any AI product. His own job is the evidence. David went from writing code effectively all day to writing very little of it, and what replaced it is planning, specification, customer conversation, and testing. Unit tests used to tell you a thing worked or it did not. Probabilistic systems do not offer that, so the work moves into evals, harnesses, and pipelines that decide what a model is allowed to do and catch it when it is wrong. He also walks through the week a feature his team had spent real time building was made redundant overnight by a vendor release, and why the right response was to stop defending it. The episode covers the dependency almost nobody puts on a slide. A company building on frontier models sits on a chain of counterparties it does not control, the same shape as the managing general agent chain from the last episode. Prices move, versions change, providers go down, and the tone of a generated document can shift while the customer assumes nothing changed. The Advocate Market Terminal is built on carrier pricing, premium, and compliance data across commercial real estate lines, and the standard is the same in both directions. If a number is not testable, it does not ship. The takeaway is three questions to ask anyone demoing AI software. How do your own engineers use it, how do you benchmark the output and show the math behind it, and how much of this rests on a single model. Learn more about what we are building: https://advocate.app/?utm_source=spotify&utm_medium=podcast #AdvocateInsuranceDesk #AdvocateTechnologies 0:00 The model is not the product 1:22 How the engineering job changed 3:47 Why model selection is overrated 5:10 What code is no longer worth writing 8:44 Hiring for problems, not for code 10:20 When a model absorbs what you built 13:12 What stays in human hands, and who is liable 15:03 Model supply chain, borrowed from MGAs 19:38 Tone drift in the proposal generator 22:58 Accuracy versus precision, and the bullseye 25:38 How to pressure test an AI vendor 28:17 A year out: the gap widens
Embed this episode
Ready to play
Building AI Software That You Can Actually Benchmark | David Haddad
No transcript for this episode yet
Similar Episodes
No similar episodes found.