EPISODE · Apr 27, 2026
Ep 623: GPT 5.5 Tops Expert Benchmarks — and Knows It — Apr 27
GPT 5.5 scores 85% on GDP-val, where evaluators prefer or tie it against human experts with 12+ years experience. Apollo Research found near-zero sandbagging, but 22% of outputs explicitly note the model knows it's being evaluated. Source episodes: - Day 1 of The 2026 AI Advantage Summit 2026 (The AI Advantage) https://www.youtube.com/watch?v=d3oI3x4mjKc - Day 3 of The 2026 AI Advantage Summit 2026 (The AI Advantage) https://www.youtube.com/watch?v=1jikcQVRVNo - OpenAI just WON... (Wes Roth) https://www.youtube.com/watch?v=evVs-Jtor50 - First impressions of GPT-5.5 from Aaron Friel (OpenAI) https://www.youtube.com/watch?v=KKiwxLK59YQ
Embed this episode
Ready to play
Ep 623: GPT 5.5 Tops Expert Benchmarks — and Knows It — Apr 27
No transcript for this episode yet
Similar Episodes
No similar episodes found.