LLM Training episode artwork

EPISODE · Jan 20, 2025 · 25 MIN

LLM Training

from Large Language Model (LLM) Talk · host AI-Talk

Training large language models (LLMs) is challenging due to the large amount of GPU memory and long training times required. Several parallelism paradigms enable model training across multiple GPUs, and various model architecture and memory-saving designs make it possible to train very large neural networks. The optimal model size and number of training tokens should be scaled equally, with a doubling of model size requiring a doubling of training tokens. Current large language models are significantly under-trained. Techniques such as data parallelism, model parallelism, pipeline parallelism, and tensor parallelism can be used to distribute the training workload. Other strategies include CPU offloading, activation recomputation, mixed-precision training, and compression to save memory.

Episode metadata supplied by the publisher feed · Published Jan 20, 2025

Embed this episode

NOW PLAYING

LLM Training

0:00 25:42

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Large Language Model (LLM) Talk?

This episode is 25 minutes long.

When was this Large Language Model (LLM) Talk episode published?

This episode was published on January 20, 2025.

Can I download this Large Language Model (LLM) Talk episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!