Skip to content
Building AI Software That You Can Actually Benchmark | David Haddad episode artwork

EPISODE · Aug 5, 2026 · 30 MIN

Building AI Software That You Can Actually Benchmark | David Haddad

from Advocate Insurance Desk · host Advocate Technologies

Every vendor in commercial insurance now says they use AI. Almost none of them will tell a buyer which part of an answer was calculated and which part was generated. David Haddad, head of product engineering at Advocate Technologies and the builder of the World Insurance Model covered in Episode 21, joins Katie Dowson to explain why that distinction is the whole thing, and why the model itself is the least interesting part of any AI product. His own job is the evidence. David went from writing code effectively all day to writing very little of it, and what replaced it is planning, specification, customer conversation, and testing. Unit tests used to tell you a thing worked or it did not. Probabilistic systems do not offer that, so the work moves into evals, harnesses, and pipelines that decide what a model is allowed to do and catch it when it is wrong. He also walks through the week a feature his team had spent real time building was made redundant overnight by a vendor release, and why the right response was to stop defending it. The episode covers the dependency almost nobody puts on a slide. A company building on frontier models sits on a chain of counterparties it does not control, the same shape as the managing general agent chain from the last episode. Prices move, versions change, providers go down, and the tone of a generated document can shift while the customer assumes nothing changed. The Advocate Market Terminal is built on carrier pricing, premium, and compliance data across commercial real estate lines, and the standard is the same in both directions. If a number is not testable, it does not ship. The takeaway is three questions to ask anyone demoing AI software. How do your own engineers use it, how do you benchmark the output and show the math behind it, and how much of this rests on a single model. Learn more about what we are building: https://advocate.app/?utm_source=spotify&utm_medium=podcast #AdvocateInsuranceDesk #AdvocateTechnologies 0:00 The model is not the product 1:22 How the engineering job changed 3:47 Why model selection is overrated 5:10 What code is no longer worth writing 8:44 Hiring for problems, not for code 10:20 When a model absorbs what you built 13:12 What stays in human hands, and who is liable 15:03 Model supply chain, borrowed from MGAs 19:38 Tone drift in the proposal generator 22:58 Accuracy versus precision, and the bullseye 25:38 How to pressure test an AI vendor 28:17 A year out: the gap widens

Episode metadata supplied by the publisher feed · Published Aug 5, 2026

Embed this episode

Ready to play

Building AI Software That You Can Actually Benchmark | David Haddad

0:00 30:25

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

How long is this episode of Advocate Insurance Desk?

This episode is 30 minutes long.

When was this Advocate Insurance Desk episode published?

This episode was published on August 5, 2026.

Can I download this Advocate Insurance Desk episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!