Is AI Benchmarking Broken? The Truth Behind "con@64" Revealed Brought to you by Avonetics.com episode artwork

EPISODE · Feb 20, 2025 · 9 MIN

Is AI Benchmarking Broken? The Truth Behind "con@64" Revealed Brought to you by Avonetics.com

from Beaker Banter · host Beaker Banter

Discover the controversial "con@64" technique, where AI models are prompted 64 times to reach a consensus answer. Is this a legitimate way to reduce variance or a sneaky trick to inflate benchmark scores? Dive into the heated debate on whether this practice skews real-world performance comparisons and unfairly impacts perceptions of model capabilities. Learn why some accuse XAI engineers of overhyping AI and how differing "con" values could be misleading the industry. For advertising opportunities, visit Avonetics.com.

Episode metadata supplied by the publisher feed · Published Feb 20, 2025

Embed this episode

NOW PLAYING

Is AI Benchmarking Broken? The Truth Behind "con@64" Revealed Brought to you by Avonetics.com

0:00 9:25

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

How long is this episode of Beaker Banter?

This episode is 9 minutes long.

When was this Beaker Banter episode published?

This episode was published on February 20, 2025.

Can I download this Beaker Banter episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!