LoRA: Low-Rank Adaptation of Large Language Models episode artwork

EPISODE · Apr 26, 2026 · 13 MIN

LoRA: Low-Rank Adaptation of Large Language Models

from Mastering Language Models: From Architecture to Optimization

The paper that made fine-tuning feel modular. Maya and Leo open up LoRA's central trick — freeze the pretrained weights and learn the update as the product of two thin matrices, around sixty-five thousand trainable numbers standing in for sixteen million — then follow it through the merge fork (flatten for zero-overhead serving, or keep adapters swappable on one frozen base), the rank and alpha dials, and why low intrinsic dimension makes the whole bet plausible. Two real practitioner arguments get staged on air: attention-only versus broad target modules, and whether LoRA's cheapness has made fine-tuning the default move when it shouldn't be. The hospital discharge-note summarizer returns to show why frozen storage is not frozen behavior. Sources: • LoRA: Low-Rank Adaptation of Large Language Models: https://arxiv.org/pdf/2106.09685

Episode metadata supplied by the publisher feed · Published Apr 26, 2026

Embed this episode

NOW PLAYING

LoRA: Low-Rank Adaptation of Large Language Models

0:00 13:38

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Mastering Language Models: From Architecture to Optimization?

This episode is 13 minutes long.

When was this Mastering Language Models: From Architecture to Optimization episode published?

This episode was published on April 26, 2026.

Can I download this Mastering Language Models: From Architecture to Optimization episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!