Episode 49: Rethinking AI Agent Evaluations episode artwork

EPISODE · Apr 17, 2026 · 30 MIN

Episode 49: Rethinking AI Agent Evaluations

from AWS re:Think Podcast

In this episode we explore how companies should evaluate AI agents across multiple dimensions — including correctness, tool selection, multi-turn reasoning, and safety . The conversation covers building reliable evaluation frameworks, balancing automated vs. human-in-the-loop testing, and leveraging observability to debug agent behavior in production.Links from the ShowAgentCore Evaluation: https://github.com/awslabs/agentcore-samples/tree/main/01-tutorials/07-AgentCore-evaluationsStrands Evaluation: https://strandsagents.com/docs/user-guide/evals-sdk/quickstart/AWS Hosts: Nolan Chen & Malini ChatterjeeEmail Your Feedback: [email protected]

Episode metadata supplied by the publisher feed · Published Apr 17, 2026

Embed this episode

AWS Sr Applied Scientist Ishan Singh joins us as we discuss how companies should evaluate AI agents across multiple dimensions

Distinct summary based on available episode metadata or transcript content.

Ready to play

Episode 49: Rethinking AI Agent Evaluations

0:00 30:08

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

How long is this episode of AWS re:Think Podcast?

This episode is 30 minutes long.

When was this AWS re:Think Podcast episode published?

This episode was published on April 17, 2026.

Can I download this AWS re:Think Podcast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!