EPISODE · Aug 18, 2023 · 33 MIN
706: Large Language Model Leaderboards and Benchmarks
from Super Data Science: ML & AI Podcast with Jon Krohn · host Jon Krohn
In this episode, Caterina Constantinescu dives deep into Large Language Models (LLMs), spotlighting top leaderboards, evaluation benchmarks, and real-world user perceptions. Plus, discover the challenges of dataset contamination and the intricacies of platforms like HELM and Chatbot Arena.Additional materials: www.superdatascience.com/706Interested in sponsoring a SuperDataScience Podcast episode? Visit JonKrohn.com/podcast for sponsorship information.
Embed this episode
NOW PLAYING
706: Large Language Model Leaderboards and Benchmarks
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.