The Evaluation Crisis episode artwork

EPISODE · Nov 26, 2025 · 15 MIN

The Evaluation Crisis

from NeurIPS 2025 by Basis Set · host Basis Set

"Passes the bar exam" doesn't mean AI can practice law. "Beats humans on ImageNet" doesn't mean it understands images. You'll learn why most AI benchmarks are fundamentally broken through the cautionary tale of the "infant morality study"—researchers thought babies preferred moral helpers, but they just liked bouncing balls. The Clever Hans effect is alive and well in 2025. If you're evaluating AI products, making purchasing decisions, or relying on AI benchmark claims, this episode gives you the critical thinking tools to cut through the nonsense. Topics Covered - Construct validity: Are we testing what we think we're testing? - The anthropomorphism trap: projecting human limitations onto AI - Why "passing the bar exam" doesn't mean AI can practice law - The Clever Hans problem in modern AI - EU AI Act and regulatory approaches - Testing AI like we test babies and animals (alien intelligence framework)

Episode metadata supplied by the publisher feed · Published Nov 26, 2025

Embed this episode

NOW PLAYING

The Evaluation Crisis

0:00 15:23

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of NeurIPS 2025 by Basis Set?

This episode is 15 minutes long.

When was this NeurIPS 2025 by Basis Set episode published?

This episode was published on November 26, 2025.

Can I download this NeurIPS 2025 by Basis Set episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!