EPISODE · Jul 27, 2026 · 57 MIN
The Model Found a Way Out - with Florian Brand (Prime Intellect)
from The Information Bottleneck · host Ravid Shwartz-Ziv & Allen Roush
Florian Brand builds evals at Prime Intellect. The premise of the conversation is that writing a benchmark is the easy part now. Keeping the model from cheating it is the job, and it takes longer than the benchmark itself.We get into why he thinks you can't evaluate a model apart from the CLI it runs in, what happens to statistics when a single run costs five figures, and whether the feeling that a model just works can ever become a number.He also has a few stories about agents finding their way around the scoring that are worth hearing cold.Timeline00:13 Intro01:00 What evals are for04:05 Agentic benchmarks07:10 Kimi K2 and model diversity08:23 Long-horizon coding tasks10:29 Building a benchmark12:15 MirrorCode14:27 Rubrics and LLM judges16:30 The cost of expert labelers17:49 Long runs and variance19:44 Evaluating the harness24:29 Chinese labs building CLIs30:00 More reward hacking37:45 Tau-bench and economic tasks39:43 Benchmaxxing and GLM 5.245:15 Statistics and cost47:56 Frontier convergence52:04 Misuse in open and closed models55:35 Self-improvementMusic"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.AboutThe Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.
Embed this episode
NOW PLAYING
The Model Found a Way Out - with Florian Brand (Prime Intellect)
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.