EPISODE · May 2, 2026
Do Language Models Know Their Limits
from AI Post Transformers
This episode explores whether large language models can genuinely recognize the limits of their own knowledge or whether they have simply learned to sound uncertain in socially acceptable ways. It examines the paper’s idea of “self-knowledge” through the lens of confidence calibration, including the dangerous case where a model does not know an answer but responds with unwarranted confidence. The discussion walks through the SelfAware benchmark, explaining how it pairs unanswerable questions with semantically similar answerable ones and why that design is both insightful and methodologically slippery. Listeners would find it interesting because it gets past simple accuracy scores and asks a more consequential question for AI safety and product reliability: when a model says “I don’t know,” is that real judgment or just polished behavior? Sources: 1. Do Large Language Models Know What They Don't Know? — Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, Xuanjing Huang, 2023 http://arxiv.org/abs/2305.18153 2. Finetuned Language Models Are Zero-Shot Learners — Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew Dai, Quoc V. Le, 2021 https://scholar.google.com/scholar?q=Finetuned+Language+Models+Are+Zero-Shot+Learners 3. Multitask Prompted Training Enables Zero-Shot Task Generalization — Victor Sanh, Albert Webson, Colin Raffel and many coauthors, 2021 https://scholar.google.com/scholar?q=Multitask+Prompted+Training+Enables+Zero-Shot+Task+Generalization 4. Training Language Models to Follow Instructions with Human Feedback — Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida and many coauthors, 2022 https://scholar.google.com/scholar?q=Training+Language+Models+to+Follow+Instructions+with+Human+Feedback 5. Scaling Instruction-Finetuned Language Models — Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, Xuezhi Wang, Denny Zhou, Quoc V. Le, Jason Wei and many coauthors, 2022 https://scholar.google.com/scholar?q=Scaling+Instruction-Finetuned+Language+Models 6. Self-Instruct: Aligning Language Models with Self-Generated Instructions — Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, Hannaneh Hajishirzi, 2023 https://scholar.google.com/scholar?q=Self-Instruct%3A+Aligning+Language+Models+with+Self-Generated+Instructions 7. Measuring and Improving Factuality in Large Language Models with Calibrated Confidence Scores — Saurav Kadavath, Eric Wallace, Luyu Gao, et al., 2022 https://scholar.google.com/scholar?q=Measuring+and+Improving+Factuality+in+Large+Language+Models+with+Calibrated+Confidence+Scores 8. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models — Aarohi Srivastava, Jos Rozen, Francesco Tintarev, et al., 2022 https://scholar.google.com/scholar?q=Beyond+the+Imitation+Game%3A+Quantifying+and+extrapolating+the+capabilities+of+language+models 9. Teaching Small Language Models to Reason — Jason Wei, Xuezhi Wang, Dale Schuurmans, et al., 2022 https://scholar.google.com/scholar?q=Teaching+Small+Language+Models+to+Reason 10. Self-Consistency Improves Chain of Thought Reasoning in Language Models — Xuezhi Wang, Jason Wei, Dale Schuurmans, et al., 2022 https://scholar.google.com/scholar?q=Self-Consistency+Improves+Chain+of+Thought+Reasoning+in+Language+Models 11. SimCSE: Simple Contrastive Learning of Sentence Embeddings — Tianyu Gao, Xingcheng Yao, Danqi Chen, 2021 https://scholar.google.com/scholar?q=SimCSE%3A+Simple+Contrastive+Learning+of+Sentence+Embeddings 12. SQuAD 2.0: The Stanford Question Answering Dataset — Pranav Rajpurkar, Robin Jia, Percy Liang, 2018 https://scholar.google.com/scholar?q=SQuAD+2.0%3A+The+Stanford+Question+Answering+Dataset 13. Uncertainty Distillation: Teaching Language Models to Express Semantic Confidence — Sophia Hager et al., 2025 https://scholar.google.com/scholar?q=Uncertainty+Distillation%3A+Teaching+Language+Models+to+Express+Semantic+Confidence 14. Large Language Model Uncertainty Measurement and Calibration for Medical Diagnosis and Treatment — Thomas Savage et al., 2024 https://scholar.google.com/scholar?q=Large+Language+Model+Uncertainty+Measurement+and+Calibration+for+Medical+Diagnosis+and+Treatment 15. Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey — Xiaoou Liu et al., 2025 https://scholar.google.com/scholar?q=Uncertainty+Quantification+and+Confidence+Calibration+in+Large+Language+Models%3A+A+Survey 16. Unanswerability Evaluation for Retrieval Augmented Generation — Xiangyu Peng, Prafulla Kumar Choubey, Caiming Xiong, Chien-Sheng Wu, 2025 https://scholar.google.com/scholar?q=Unanswerability+Evaluation+for+Retrieval+Augmented+Generation 17. Answerability in Retrieval-Augmented Open-Domain Question Answering — Rustam Abdumalikov, Pasquale Minervini, Yova Kementchedjhieva, 2024 https://scholar.google.com/scholar?q=Answerability+in+Retrieval-Augmented+Open-Domain+Question+Answering 18. The Art of Saying No: Contextual Noncompliance in Language Models — Faeze Brahman et al., 2024 https://scholar.google.com/scholar?q=The+Art+of+Saying+No%3A+Contextual+Noncompliance+in+Language+Models 19. AI Post Transformers: Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Model — Hal Turing & Dr. Ada Shannon, 2025 https://podcast.do-not-panic.com/episodes/hallucination-to-truth-a-review-of-fact-checking-and-factuality-evaluation-in-la/ 20. AI Post Transformers: Internal Safety Collapse in Frontier LLMs — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-04-internal-safety-collapse-in-frontier-llm-8be72f.mp3 21. AI Post Transformers: Self-Search Reinforcement Learning for LLMs — Hal Turing & Dr. Ada Shannon, 2025 https://podcast.do-not-panic.com/episodes/self-search-reinforcement-learning-for-llms/ 22. AI Post Transformers: Real Context Size and Context Rot — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-07-real-context-size-and-context-rot-56cbb4.mp3 Interactive Visualization: Do Language Models Know Their Limits
Embed this episode
NOW PLAYING
Do Language Models Know Their Limits
No transcript for this episode yet
Similar Episodes
No similar episodes found.