Rethinking Uncertainty Evaluation in Large Language Models
What we actually need is for LLM confidence estimates to satisfy the conditions required of coherent probabilistic beliefs. We formalize these conditions along three axes (structural coherence, faithfulness, and usefulness) and operationalize them as the C1 metrics. RLHF and chain-of-thought improve usefulness metrics without restoring coherence.
- ▪What we actually need is for LLM confidence estimates to satisfy the conditions required of coherent probabilistic beliefs.
- ▪We formalize these conditions along three axes (structural coherence, faithfulness, and usefulness) and operationalize them as the C1 metrics.
- ▪RLHF and chain-of-thought improve usefulness metrics without restoring coherence.
Opening excerpt (first ~120 words) tap to expand
Computer Science > Artificial Intelligence arXiv:2607.19367 (cs) [Submitted on 11 Jun 2026] Title:Rethinking Uncertainty Evaluation in Large Language Models Authors:Krish Matta, Atharv Naphade, Andy Zou View a PDF of the paper titled Rethinking Uncertainty Evaluation in Large Language Models, by Krish Matta and 2 other authors View PDF HTML (experimental) Abstract:Calibration is the primary criterion for evaluating LLM confidence, but it is insufficient: it admits trivially incoherent estimators, depends on the evaluation distribution, and does not test the extent to which the estimation can be interpreted as a consistent, underlying probability function. What we actually need is for LLM confidence estimates to satisfy the conditions required of coherent probabilistic beliefs.
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at arXiv.org.