LLM-as-a-judge: How Shopify built AI judges that match human reviewers | Spencer Lawrence episode artwork

EPISODE · Apr 17, 2025 · 45 MIN

LLM-as-a-judge: How Shopify built AI judges that match human reviewers | Spencer Lawrence

from AI Adoption Playbook · host Credal

What happens when AI capabilities outpace organizational readiness? At Shopify, this tension has pushed them to develop a practical implementation approach that balances rapid experimentation with sustainable value creation.  Spencer Lawrence, Director of Data Science & Engineering, shares how they've evolved from simple text expansion experiments to sophisticated AI assistants like Help Center and Sidekick that are transforming both customer support and merchant operations. At the heart of their strategy is a barbell approach enabling self-service for small AI use cases while making targeted investments in transformative projects. Spencer also explains how their one-week sprint cycles, sophisticated evaluation frameworks, and cross-functional collaboration have helped them overcome the common challenges that prevent organizations from realizing AI's full potential. Successful AI implementation requires more than just technical solutions — it demands new organizational structures, evaluation methods, and a willingness to constantly reevaluate what knowledge work means in an AI-augmented world. Topics discussed: Shopify's evolution from early text expansion experiments to production-level AI assistants that support both customers and merchants. Creating sophisticated evaluation frameworks that combine human annotators with LLM judges to ensure quality and consistency of AI outputs. Implementing a barbell strategy that balances small self-service AI use cases with strategic investments in high-impact projects. Running one-week sprints across all AI work to maximize iteration cycles and maintain velocity even at enterprise scale. Addressing the gap between AI capabilities and real-world impact through both technological solutions and organizational change. Building feedback loops between technical teams and legal/compliance departments to create AI solutions that meet governance requirements. Fostering a culture that values experimentation while developing clear policies that give employees confidence to innovate responsibly. Exploring how AI will raise productivity expectations rather than simply reducing workloads across all roles and functions. Using AI as a strategic thought partner to generate novel ideas and help evaluate different perspectives on complex problems. Developing a forward-looking perspective on knowledge work that embraces AI augmentation while maintaining human judgment and oversight. Listen to more episodes:  Apple  Spotify  YouTube

Episode metadata supplied by the publisher feed · Published Apr 17, 2025

Embed this episode

Ready to play

LLM-as-a-judge: How Shopify built AI judges that match human reviewers | Spencer Lawrence

0:00 45:35

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

How long is this episode of AI Adoption Playbook?

This episode is 45 minutes long.

When was this AI Adoption Playbook episode published?

This episode was published on April 17, 2025.

Can I download this AI Adoption Playbook episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!