EPISODE · Jul 1, 2026
Scaling Prompt Tuning for Frozen T5 Models
from AI Post Transformers
This episode explores Brian Lester et al.’s 2021 paper on prompt tuning, which asks whether a large frozen T5 model can be adapted to new tasks by learning only a tiny soft prompt instead of fine-tuning all model weights. It explains the difference between soft prompt tuning, full fine-tuning, prefix-tuning, and GPT-3-style few-shot prompting, and frames the paper as a test of whether scaling laws make lightweight adaptation dramatically more effective at large model sizes. The discussion highlights the key result that prompt tuning lags on smaller models but approaches full fine-tuning on very large T5 checkpoints, with longer prompts and vocabulary-based initialization helping, while a five-token prompt can shrink task-specific parameters from 11 billion to roughly 20,000. Listeners would find it interesting because it connects model-scaling theory to concrete engineering tradeoffs around storage, mixed-task serving, and why industry later gravitated toward PEFT methods like LoRA and adapters. Sources: 1. The Power of Scale for Parameter-Efficient Prompt Tuning — Brian Lester, Rami Al-Rfou, Noah Constant, 2021 http://arxiv.org/abs/2104.08691 2. Prefix-Tuning: Optimizing Continuous Prompts for Generation — Xiang Lisa Li, Percy Liang, 2021 https://arxiv.org/abs/2101.00190 3. Learning How to Ask: Querying LMs with Mixtures of Soft Prompts — Guanghui Qin, Jason Eisner, 2021 https://arxiv.org/abs/2104.06599 4. The Power of Scale for Parameter-Efficient Prompt Tuning — Brian Lester, Rami Al-Rfou, Noah Constant, 2021 https://arxiv.org/abs/2104.08691 5. Multitask Prompt Tuning Enables Parameter-Efficient Transfer Learning — Zhen Wang, Rameswar Panda, Leonid Karlinsky, Rogerio Feris, Huan Sun, Yoon Kim, 2023 https://arxiv.org/abs/2303.02861 6. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer — Colin Raffel et al., 2020 https://scholar.google.com/scholar?q=Exploring+the+Limits+of+Transfer+Learning+with+a+Unified+Text-to-Text+Transformer 7. Language Models are Few-Shot Learners — Tom B. Brown et al., 2020 https://scholar.google.com/scholar?q=Language+Models+are+Few-Shot+Learners 8. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts — Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, Sameer Singh, 2020 https://scholar.google.com/scholar?q=AutoPrompt%3A+Eliciting+Knowledge+from+Language+Models+with+Automatically+Generated+Prompts 9. WARP: Word-level Adversarial ReProgramming — Karen Hambardzumyan, Hrant Khachatrian, Jonathan May, 2021 https://scholar.google.com/scholar?q=WARP%3A+Word-level+Adversarial+ReProgramming 10. Parameter-Efficient Transfer Learning for NLP — Neil Houlsby et al., 2019 https://scholar.google.com/scholar?q=Parameter-Efficient+Transfer+Learning+for+NLP 11. MRQA 2019 Shared Task: Evaluating Generalization in Reading Comprehension — Adam Fisch, Alon Talmor, Robin Jia, Minjoon Seo, Eunsol Choi, Danqi Chen, 2019 https://scholar.google.com/scholar?q=MRQA+2019+Shared+Task%3A+Evaluating+Generalization+in+Reading+Comprehension 12. Task Prompt Vectors: Effective Initialization through Multi-Task Soft-Prompt Transfer — Robert Belanec, Simon Ostermann, Ivan Srba, Maria Bielikova, 2024 https://scholar.google.com/scholar?q=Task+Prompt+Vectors%3A+Effective+Initialization+through+Multi-Task+Soft-Prompt+Transfer 13. Parameter-Efficient Fine-Tuning for Medical Text Summarization: A Comparative Study of LoRA, Prompt Tuning, and Full Fine-Tuning — Ulugbek Shernazarov et al., 2026 https://scholar.google.com/scholar?q=Parameter-Efficient+Fine-Tuning+for+Medical+Text+Summarization%3A+A+Comparative+Study+of+LoRA%2C+Prompt+Tuning%2C+and+Full+Fine-Tuning 14. MerA: Merging Pretrained Adapters For Few-Shot Learning — Shwai He et al., 2023 https://scholar.google.com/scholar?q=MerA%3A+Merging+Pretrained+Adapters+For+Few-Shot+Learning 15. Exploring the Relationship between In-Context Learning and Instruction Tuning — Hanyu Duan et al., 2023 https://scholar.google.com/scholar?q=Exploring+the+Relationship+between+In-Context+Learning+and+Instruction+Tuning 16. Is In-Context Learning Sufficient for Instruction Following in LLMs? — Hao Zhao et al., 2024 https://scholar.google.com/scholar?q=Is+In-Context+Learning+Sufficient+for+Instruction+Following+in+LLMs%3F 17. Symbol tuning improves in-context learning in language models — Jerry Wei et al., 2023 https://scholar.google.com/scholar?q=Symbol+tuning+improves+in-context+learning+in+language+models 18. Last One Standing: A Comparative Analysis of Security and Privacy of Soft Prompt Tuning, LoRA, and In-Context Learning — Rui Wen et al., 2023 https://scholar.google.com/scholar?q=Last+One+Standing%3A+A+Comparative+Analysis+of+Security+and+Privacy+of+Soft+Prompt+Tuning%2C+LoRA%2C+and+In-Context+Learning 19. AI Post Transformers: Benchmarking PEFT Techniques for Large Language Models — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-06-20-benchmarking-peft-techniques-for-large-l-41bbf5.mp3
Embed this episode
NOW PLAYING
Scaling Prompt Tuning for Frozen T5 Models
No transcript for this episode yet
Similar Episodes
No similar episodes found.