EPISODE · Jun 16, 2026 · 22 MIN
Why Your AI Focus Group Keeps Saying Three
from The Digital Transformation Playbook · host Kieran Gilmurray
You spend years building a product, polish the packaging, nail the pitch… then you hit the terrifying question: is anyone actually going to buy it? We dig into a 2025 research result from PyMC Labs and Colgate-Palmolive that aims straight at that fear with AI market research, synthetic consumers, and large language models that can simulate purchase intent at scale.TL;DR / At A Glancethe core problem with direct Likert ratings and why LLMs collapse to neutral threeshow semantic similarity rating converts free-text responses into numerical scores using embeddings and cosine similaritywhy follow-up AI grading helps but still trails the embedding-based approachwhat 57 real product surveys and 9,300 human responses reveal about accuracy and distribution matchinghow persona prompting reproduces real demographic patterns across age and income constraintswhy zero-shot LLM methods can beat supervised machine learning models trained on the same domainThe shocker is that the first attempt fails badly. When you make models like GPT-4 or Gemini answer a classic Likert scale with a single number, they hedge and pile up on neutral “3” ratings. The fix is not “better AI”, it is better questioning. Google Notebook LM Agents help us unpack semantic similarity rating: let the model respond in natural language, convert that text into embeddings, and map it to five anchor statements using cosine similarity. You get fast, automated scoring without stripping away the model’s reasoning.From there, we pressure-test the method against thousands of real survey responses across dozens of personal care product concepts, then look at whether AI personas actually reflect real constraints like age and income. We also compare the approach with traditional machine learning models such as LightGBM, and dig into an underrated advantage: synthetic consumers can produce richer, more candid qualitative feedback than many human panels.If you care about product testing, consumer insights, or the future of focus groups, listen through and tell us where you’d trust this and where you wouldn’t. Subscribe, share with a colleague, and leave a review with your take: would you let synthetic consumers influence a real launch?Paper: http://arxiv.org/abs/2510.08338Support the showIf you are leading your businesses strategic transformation and need greater clarity, stronger execution and measurable results, let’s connect. 🌎 Website: www.KieranGilmurray.com📅 Book a call: https://calendly.com/kierangilmurray/catch-up📘 Kieran Gilmurray | LinkedIn🌐 Substack: https://kierangilmurray.substack.com📕 Amazon https://tinyurl.com/MyBooksOnAmazonUK AI Transparency Notice: This podcast uses a hybrid format. When an episode features one of Kieran Gilmurray’s written articles, the narration is generated using a synthetic clone of his voice via ElevenLabs AI (the underlying article text is entirely human-authored). Episode descriptions and summaries are assisted by AI and should be considered unedited by a human unless specified.
Embed this episode
What this episode covers
You spend years building a product, polish the packaging, nail the pitch… then you hit the terrifying question: is anyone actually going to buy it? We dig into a 2025 research result from PyMC Labs and Colgate-Palmolive that aims straight at that fear with AI market research, synthetic consumers, and large language models that can simulate purchase intent at scale. TL;DR / At A Glance the core problem with direct Likert ratings and why LLMs collapse to neutral threeshow semantic similarity ra...
NOW PLAYING
Why Your AI Focus Group Keeps Saying Three
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.