EPISODE · Apr 17, 2026 · 30 MIN
Episode 49: Rethinking AI Agent Evaluations
from AWS re:Think Podcast
In this episode we explore how companies should evaluate AI agents across multiple dimensions — including correctness, tool selection, multi-turn reasoning, and safety . The conversation covers building reliable evaluation frameworks, balancing automated vs. human-in-the-loop testing, and leveraging observability to debug agent behavior in production.Links from the ShowAgentCore Evaluation: https://github.com/awslabs/agentcore-samples/tree/main/01-tutorials/07-AgentCore-evaluationsStrands Evaluation: https://strandsagents.com/docs/user-guide/evals-sdk/quickstart/AWS Hosts: Nolan Chen & Malini ChatterjeeEmail Your Feedback: [email protected]
Embed this episode
What this episode covers
AWS Sr Applied Scientist Ishan Singh joins us as we discuss how companies should evaluate AI agents across multiple dimensions
Ready to play
Episode 49: Rethinking AI Agent Evaluations
No transcript for this episode yet
Similar Episodes
No similar episodes found.