EPISODE · May 27, 2025 · 13 MIN
Your Pre-trained LLM is Secretly an Unsupervised Confidence Calibrator
from Best AI papers explained · host Enoch H. Kang
This paper introduces Disagreement-Aware Confidence Alignment (DACA), an unsupervised method for calibrating the confidence of post-trained large language models (PoLMs). While pre-trained language models (PLMs) are typically well-calibrated, post-training can lead to over-confidence, especially with limited labeled data. DACA addresses this by leveraging the well-calibrated confidence of PLMs on unlabeled data, specifically by optimizing calibration parameters only on examples where PLM and PoLM predictions agree. This process avoids the negative impact of prediction disagreement on calibration, resulting in more accurate confidence scores for PoLMs, which is shown to improve performance on various benchmarks and model sizes, including for open-ended question answering and selective classification.
Embed this episode
NOW PLAYING
Your Pre-trained LLM is Secretly an Unsupervised Confidence Calibrator
No transcript for this episode yet
Similar Episodes
No similar episodes found.