PALOMA: Benchmarking Language Model Fit Across Domains episode artwork

EPISODE · Jun 24, 2026

PALOMA: Benchmarking Language Model Fit Across Domains

from AI Post Transformers

This episode explores PALOMA, a NeurIPS 2024 benchmark designed to measure how well language models fit many different language distributions instead of relying on a single average perplexity score. It explains why one global loss number can hide important weaknesses across domains such as specific subreddits, scientific writing, or programming languages, and highlights PALOMA’s fine-grained setup across 546 English and code domains from 16 sources. The discussion places PALOMA in context with earlier language-model evaluation traditions, scaling-law work, and broader benchmark efforts like HELM, while arguing that evaluation design determines what claims researchers can actually make. Listeners would find it interesting for its clear case that better measurement, data curation, and decontamination can reveal model behavior that broad headline metrics often miss. Sources: 1. Paloma: A Benchmark for Evaluating Language Model Fit — Ian Magnusson, Akshita Bhagia, Valentin Hofmann, Luca Soldaini, Ananya Harsh Jha, Oyvind Tafjord, Dustin Schwenk, Evan Pete Walsh, Yanai Elazar, Kyle Lo, Dirk Groeneveld, Iz Beltagy, Hannaneh Hajishirzi, Noah A. Smith, Kyle Richardson, Jesse Dodge, 2023 http://arxiv.org/abs/2312.10523 2. One Billion Word Benchmark for Measuring Progress in Statistical Language Modeling — Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, et al., 2013 https://scholar.google.com/scholar?q=One+Billion+Word+Benchmark+for+Measuring+Progress+in+Statistical+Language+Modeling 3. Scaling Laws for Neural Language Models — Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Dario Amodei, et al., 2020 https://scholar.google.com/scholar?q=Scaling+Laws+for+Neural+Language+Models 4. Training Compute-Optimal Large Language Models — Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Jack W. Rae, Oriol Vinyals, Laurent Sifre, et al., 2022 https://scholar.google.com/scholar?q=Training+Compute-Optimal+Large+Language+Models 5. Paloma: A Benchmark for Evaluating Language Model Fit — Ian Magnusson, Akshita Bhagia, Valentin Hofmann, Luca Soldaini, Kyle Richardson, Jesse Dodge, et al., 2024 https://scholar.google.com/scholar?q=Paloma%3A+A+Benchmark+for+Evaluating+Language+Model+Fit 6. M2D2: A Massively Multi-domain Language Modeling Dataset — Machel Reid, Victor Zhong, Suchin Gururangan, Luke Zettlemoyer, 2022 https://scholar.google.com/scholar?q=M2D2%3A+A+Massively+Multi-domain+Language+Modeling+Dataset 7. Holistic Evaluation of Language Models — Percy Liang et al., 2022 https://scholar.google.com/scholar?q=Holistic+Evaluation+of+Language+Models 8. Same Pre-training Loss, Better Downstream: Implicit Bias Matters for Language Models — Hong Liu, Sang Michael Xie, Zhiyuan Li, Tengyu Ma, 2022 https://scholar.google.com/scholar?q=Same+Pre-training+Loss%2C+Better+Downstream%3A+Implicit+Bias+Matters+for+Language+Models 9. Language Model Evaluation Beyond Perplexity — Clara Meister, Ryan Cotterell, 2021 https://scholar.google.com/scholar?q=Language+Model+Evaluation+Beyond+Perplexity 10. Unsupervised Domain Clusters in Pretrained Language Models — Roee Aharoni, Yoav Goldberg, 2020 https://scholar.google.com/scholar?q=Unsupervised+Domain+Clusters+in+Pretrained+Language+Models 11. DataComp-LM: In search of the next generation of training sets for language models — Jeffrey Li et al., 2024 https://scholar.google.com/scholar?q=DataComp-LM%3A+In+search+of+the+next+generation+of+training+sets+for+language+models 12. Rethinking Perplexity: Revealing the Impact of Input Length on Perplexity Evaluation in LLMs — Letian Cheng et al., 2026 https://scholar.google.com/scholar?q=Rethinking+Perplexity%3A+Revealing+the+Impact+of+Input+Length+on+Perplexity+Evaluation+in+LLMs 13. Rethinking Benchmark and Contamination for Language Models with Rephrased Samples — Shuo Yang et al., 2023 https://scholar.google.com/scholar?q=Rethinking+Benchmark+and+Contamination+for+Language+Models+with+Rephrased+Samples 14. PaCoST: Paired Confidence Significance Testing for Benchmark Contamination Detection in Large Language Models — Huixuan Zhang et al., 2024 https://scholar.google.com/scholar?q=PaCoST%3A+Paired+Confidence+Significance+Testing+for+Benchmark+Contamination+Detection+in+Large+Language+Models 15. RegMix: Data Mixture as Regression for Language Model Pre-training — Qian Liu et al., 2024 https://scholar.google.com/scholar?q=RegMix%3A+Data+Mixture+as+Regression+for+Language+Model+Pre-training 16. Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance — Jiasheng Ye et al., 2024 https://scholar.google.com/scholar?q=Data+Mixing+Laws%3A+Optimizing+Data+Mixtures+by+Predicting+Language+Modeling+Performance 17. The Fools are Certain; the Wise are Doubtful: Exploring LLM Confidence in Code Completion — Zoe Kotti et al., 2025 https://scholar.google.com/scholar?q=The+Fools+are+Certain%3B+the+Wise+are+Doubtful%3A+Exploring+LLM+Confidence+in+Code+Completion 18. AI Post Transformers: Model-Aware Tokenizer Transfer for Multilingual LLMs — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-03-16-model-aware-tokenizer-transfer-for-multi-90666c.mp3 19. AI Post Transformers: LLM Benchmark Robustness to Linguistic Variation — Hal Turing & Dr. Ada Shannon, 2025 https://podcast.do-not-panic.com/episodes/llm-benchmark-robustness-to-linguistic-variation/

Episode metadata supplied by the publisher feed · Published Jun 24, 2026

Embed this episode

NOW PLAYING

PALOMA: Benchmarking Language Model Fit Across Domains

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on June 24, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!