o3 - wow episode artwork

EPISODE · Dec 21, 2024 · 22 MIN

o3 - wow

from AI Explained Official Podcast · host Philip - Host of AI Explained YT

o3 isn’t one of the biggest developments in AI for 2+ years because it beats a particular benchmark. It is so because it demonstrates a reusable technique through which almost any benchmark could fall, and at short notice. I’ll cover all the highlights, benchmarks broken, and what comes next. Plus, the costs OpenAI didn’t want us to know, Genesis, ARC-AGI 2, Gemini-Thinking, and much more. FrontierMath: https://epoch.ai/frontiermathhttps://arxiv.org/pdf/2411.04872Chollet Statement:https://arcprize.org/blog/oai-o3-pub-breakthroughMLC Paper: https://www.scientificamerican.com/article/new-training-method-helps-ai-generalize-like-people-do/?utm_campaign=socialflow&utm_source=twitter&utm_medium=socialAlphaCode 2: https://storage.googleapis.com/deepmind-media/AlphaCode2/AlphaCode2_Tech_Report.pdfHuman Performance on ARC-AGI: https://arxiv.org/pdf/2409.01374v1Wei Tweet ‘3 months’:https://x.com/_jasonwei/status/1870184982007644614Deliberative Alignment Paper: https://openai.com/index/deliberative-alignment/Brown Safety Tweet: https://x.com/polynoamial/status/1870196476908834893Swe-Bench Verified: https://openai.com/index/introducing-swe-bench-verified/Amodei Prediction: https://x.com/OfirPress/status/1858567863788769518David Dohan: 16 hours https://x.com/dmdohan/status/1870171404093796638OpenAI Personal Writing: https://openai.com/index/learning-to-reason-with-llms/https://simple-bench.com/John Hallman Tweet: https://x.com/johnohallman/status/187023337568194572500:00 - Introduction01:19 - What is o3?03:18 - FrontierMath05:15 - o4, o506:03 - GPQA06:24 - Coding, Codeforces + SWE-verified, AlphaCode 208:13 - 1st Caveat09:03 - Compositionality?10:16 - SimpleBench?13:11 - ARC-AGI, Chollet

Episode metadata supplied by the publisher feed · Published Dec 21, 2024

Embed this episode

o3 isn’t one of the biggest developments in AI for 2+ years because it beats a particular benchmark. It is so because it demonstrates a reusable technique through which almost any benchmark could fall, and at short notice. I’ll cover all the highlights, benchmarks broken, and what comes next. Plus, the costs OpenAI didn’t want us to know, Genesis, ARC-AGI 2, Gemini-Thinking, and much more. FrontierMath: https://epoch.ai/frontiermath https://arxiv.org/pdf/2411.04872 Chollet Statement:h...

Distinct summary based on available episode metadata or transcript content.

Ready to play

o3 - wow

0:00 22:20

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of AI Explained Official Podcast?

This episode is 22 minutes long.

When was this AI Explained Official Podcast episode published?

This episode was published on December 21, 2024.

Can I download this AI Explained Official Podcast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!