Improving the Calibration of Confidence Scores in Text Generation Using the Output Distribution's Characteristics
Fuente:
arXiv
Saved in:
| Main Authors: | Flores, Lorenzo Jaime Yu, Ernst, Ori, Cheung, Jackie Chi Kit |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PreSumm: Predicting Summarization Performance Without Summarizing
by: Koniaev, Steven, et al.
Published: (2025)
by: Koniaev, Steven, et al.
Published: (2025)
Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2026)
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2026)
Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations
by: Yu, Lei, et al.
Published: (2024)
by: Yu, Lei, et al.
Published: (2024)
$\texttt{COSMIC}$: Mutual Information for Task-Agnostic Summarization Evaluation
by: Darrin, Maxime, et al.
Published: (2024)
by: Darrin, Maxime, et al.
Published: (2024)
Testing the Assumptions of Active Learning for Translation Tasks with Few Samples
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2026)
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2026)
$(RSA)^2$: A Rhetorical-Strategy-Aware Rational Speech Act Framework for Figurative Language Understanding
by: Piano, Cesare Spinoso-Di, et al.
Published: (2025)
by: Piano, Cesare Spinoso-Di, et al.
Published: (2025)
Can Vision Language Models Be Adaptive in Mathematics Education? A Learner Model-based Rubric Study
by: Gao, Jie, et al.
Published: (2026)
by: Gao, Jie, et al.
Published: (2026)
Error Diversity Matters: An Error-Resistant Ensemble Method for Unsupervised Dependency Parsing
by: Shayegh, Behzad, et al.
Published: (2024)
by: Shayegh, Behzad, et al.
Published: (2024)
Calibrating Verbalized Confidence with Self-Generated Distractors
by: Wang, Victor, et al.
Published: (2025)
by: Wang, Victor, et al.
Published: (2025)
Agentic Confidence Calibration
by: Zhang, Jiaxin, et al.
Published: (2026)
by: Zhang, Jiaxin, et al.
Published: (2026)
Unsupervised Layer-wise Score Aggregation for Textual OOD Detection
by: Darrin, Maxime, et al.
Published: (2023)
by: Darrin, Maxime, et al.
Published: (2023)
From Thinking to Output: Chain-of-Thought and Text Generation Characteristics in Reasoning Language Models
by: Liu, Junhao, et al.
Published: (2025)
by: Liu, Junhao, et al.
Published: (2025)
Contextualized Sequence Likelihood: Enhanced Confidence Scores for Natural Language Generation
by: Lin, Zhen, et al.
Published: (2024)
by: Lin, Zhen, et al.
Published: (2024)
On the Benefits of Fine-Grained Loss Truncation: A Case Study on Factuality in Summarization
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2024)
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2024)
Fact-Level Confidence Calibration and Self-Correction
by: Yuan, Yige, et al.
Published: (2024)
by: Yuan, Yige, et al.
Published: (2024)
Double-Calibration: Towards Reliable LLMs via Calibrating Knowledge and Reasoning Confidence
by: Lu, Yuyin, et al.
Published: (2026)
by: Lu, Yuyin, et al.
Published: (2026)
Solving the Challenge Set without Solving the Task: On Winograd Schemas as a Test of Pronominal Coreference Resolution
by: Porada, Ian, et al.
Published: (2024)
by: Porada, Ian, et al.
Published: (2024)
Towards Trustworthy Report Generation: A Deep Research Agent with Progressive Confidence Estimation and Calibration
by: Yuan, Yi, et al.
Published: (2026)
by: Yuan, Yi, et al.
Published: (2026)
STED and Consistency Scoring: A Framework for Evaluating LLM Structured Output Reliability
by: Wang, Guanghui, et al.
Published: (2025)
by: Wang, Guanghui, et al.
Published: (2025)
Confident RAG: Enhancing the Performance of LLMs for Mathematics Question Answering through Multi-Embedding and Confidence Scoring
by: Chen, Shiting, et al.
Published: (2025)
by: Chen, Shiting, et al.
Published: (2025)
Don't Think Twice! Over-Reasoning Impairs Confidence Calibration
by: Lacombe, Romain, et al.
Published: (2025)
by: Lacombe, Romain, et al.
Published: (2025)
A Survey of Confidence Estimation and Calibration in Large Language Models
by: Geng, Jiahui, et al.
Published: (2023)
by: Geng, Jiahui, et al.
Published: (2023)
How Teachers Can Use Large Language Models and Bloom's Taxonomy to Create Educational Quizzes
by: Elkins, Sabina, et al.
Published: (2024)
by: Elkins, Sabina, et al.
Published: (2024)
LACIE: Listener-Aware Finetuning for Confidence Calibration in Large Language Models
by: Stengel-Eskin, Elias, et al.
Published: (2024)
by: Stengel-Eskin, Elias, et al.
Published: (2024)
Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals
by: Torrielli, Federico, et al.
Published: (2026)
by: Torrielli, Federico, et al.
Published: (2026)
Confidence Improves Self-Consistency in LLMs
by: Taubenfeld, Amir, et al.
Published: (2025)
by: Taubenfeld, Amir, et al.
Published: (2025)
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
by: Balkır, Esma, et al.
Published: (2026)
by: Balkır, Esma, et al.
Published: (2026)
Linguistic Characteristics of AI-Generated Text: A Survey
by: Terčon, Luka, et al.
Published: (2025)
by: Terčon, Luka, et al.
Published: (2025)
Mind the Confidence Gap: Overconfidence, Calibration, and Distractor Effects in Large Language Models
by: Chhikara, Prateek
Published: (2025)
by: Chhikara, Prateek
Published: (2025)
The Dunning-Kruger Effect in Large Language Models: An Empirical Study of Confidence Calibration
by: Ghosh, Sudipta, et al.
Published: (2026)
by: Ghosh, Sudipta, et al.
Published: (2026)
Confidence-Calibrated Small-Large Language Model Collaboration for Cost-Efficient Reasoning
by: Zhang, Chuang, et al.
Published: (2026)
by: Zhang, Chuang, et al.
Published: (2026)
Making Retrieval-Augmented Language Models Robust to Irrelevant Context
by: Yoran, Ori, et al.
Published: (2023)
by: Yoran, Ori, et al.
Published: (2023)
DecMetrics: Structured Claim Decomposition Scoring for Factually Consistent LLM Outputs
by: Huang, Minghui
Published: (2025)
by: Huang, Minghui
Published: (2025)
Calibrating Verbalized Probabilities for Large Language Models
by: Wang, Cheng, et al.
Published: (2024)
by: Wang, Cheng, et al.
Published: (2024)
MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
by: Subramani, Nishant, et al.
Published: (2025)
by: Subramani, Nishant, et al.
Published: (2025)
Detoxification of Large Language Models through Output-layer Fusion with a Calibration Model
by: Tian, Yuanhe, et al.
Published: (2025)
by: Tian, Yuanhe, et al.
Published: (2025)
Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces
by: Pathak, Manas, et al.
Published: (2026)
by: Pathak, Manas, et al.
Published: (2026)
DivScore: Zero-Shot Detection of LLM-Generated Text in Specialized Domains
by: Chen, Zhihui, et al.
Published: (2025)
by: Chen, Zhihui, et al.
Published: (2025)
OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation
by: Lage, Lucas Fonseca, et al.
Published: (2025)
by: Lage, Lucas Fonseca, et al.
Published: (2025)
Generative Ontology: When Structured Knowledge Learns to Create
by: Cheung, Benny
Published: (2026)
by: Cheung, Benny
Published: (2026)
Similar Items
-
PreSumm: Predicting Summarization Performance Without Summarizing
by: Koniaev, Steven, et al.
Published: (2025) -
Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2026) -
Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations
by: Yu, Lei, et al.
Published: (2024) -
$\texttt{COSMIC}$: Mutual Information for Task-Agnostic Summarization Evaluation
by: Darrin, Maxime, et al.
Published: (2024) -
Testing the Assumptions of Active Learning for Translation Tasks with Few Samples
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2026)