Data Mixture Optimization: A Multi-fidelity Multi-scale Bayesian Framework episode artwork

EPISODE · Mar 31, 2025 · 22 MIN

Data Mixture Optimization: A Multi-fidelity Multi-scale Bayesian Framework

from Best AI papers explained · host Enoch H. Kang

This paper  outlines a new probabilistic framework called multi-fidelity multi-scale Bayesian optimization for efficiently determining the best combinations of data sources for pre-training large language models. It addresses the limitations of intuition-based and deterministic extrapolation methods by modeling uncertainty and sequentially selecting data mixtures, model sizes, and training steps to balance cost and information gain. The authors introduce a simulator based on numerous pre-training runs to demonstrate the effectiveness of their approach, showing significant speedups compared to existing techniques. Ultimately, the work proposes a more principled and transferable method for optimizing data mixtures, acknowledging the value of information from smaller-scale experiments.

Episode metadata supplied by the publisher feed · Published Mar 31, 2025

Embed this episode

NOW PLAYING

Data Mixture Optimization: A Multi-fidelity Multi-scale Bayesian Framework

0:00 22:01

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

How long is this episode of Best AI papers explained?

This episode is 22 minutes long.

When was this Best AI papers explained episode published?

This episode was published on March 31, 2025.

Can I download this Best AI papers explained episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!