LLMs are Overconfident: Evaluating Confidence Interval Calibration with FermiEval
Fuente:
arXiv
Saved in:
| Main Authors: | Epstein, Elliot L., Winnicki, John, Sornwanee, Thanawat, Dwaraknath, Rajat |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Allocate Marginal Reviews to Borderline Papers Using LLM Comparative Ranking
by: Epstein, Elliot L., et al.
Published: (2026)
by: Epstein, Elliot L., et al.
Published: (2026)
Do Language Models Mirror Human Confidence? Exploring Psychological Insights to Address Overconfidence in LLMs
by: Xu, Chenjun, et al.
Published: (2025)
by: Xu, Chenjun, et al.
Published: (2025)
SD-KDE: Score-Debiased Kernel Density Estimation
by: Epstein, Elliot L., et al.
Published: (2025)
by: Epstein, Elliot L., et al.
Published: (2025)
Flash-SD-KDE: Accelerating SD-KDE with Tensor Cores
by: Epstein, Elliot L., et al.
Published: (2026)
by: Epstein, Elliot L., et al.
Published: (2026)
The Metacognitive Probe: Five Behavioural Calibration Diagnostics for LLMs
by: Oliveira, Rafael C. T.
Published: (2026)
by: Oliveira, Rafael C. T.
Published: (2026)
A Confidence-Diversity Framework for Calibrating AI Judgement in Accessible Qualitative Coding Tasks
by: Zhao, Zhilong, et al.
Published: (2025)
by: Zhao, Zhilong, et al.
Published: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
Evaluating Relational Reasoning in LLMs with REL
by: Fesser, Lukas, et al.
Published: (2026)
by: Fesser, Lukas, et al.
Published: (2026)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
Extreme AutoML: Analysis of Classification, Regression, and NLP Performance
by: Ratner, Edward, et al.
Published: (2024)
by: Ratner, Edward, et al.
Published: (2024)
Annif at the GermEval-2025 LLMs4Subjects Task: Traditional XMTC Augmented by Efficient LLMs
by: Suominen, Osma, et al.
Published: (2025)
by: Suominen, Osma, et al.
Published: (2025)
Evaluating Long Range Dependency Handling in Code Generation LLMs
by: Assogba, Yannick, et al.
Published: (2024)
by: Assogba, Yannick, et al.
Published: (2024)
ALBA: A European Portuguese Benchmark for Evaluating Language and Linguistic Dimensions in Generative LLMs
by: Vieira, Inês, et al.
Published: (2026)
by: Vieira, Inês, et al.
Published: (2026)
FormationEval, an open multiple-choice benchmark for petroleum geoscience
by: Ermilov, Almaz
Published: (2026)
by: Ermilov, Almaz
Published: (2026)
PSK at SemEval-2026 Task 9: Multilingual Polarization Detection Using Ensemble Gemma Models with Synthetic Data Augmentation
by: Pulipaka, Srikar Kashyap
Published: (2026)
by: Pulipaka, Srikar Kashyap
Published: (2026)
Overclocking LLM Reasoning: Monitoring and Controlling Thinking Path Lengths in LLMs
by: Eisenstadt, Roy, et al.
Published: (2025)
by: Eisenstadt, Roy, et al.
Published: (2025)
When Words Change the Model: Sensitivity of LLMs for Constraint Programming Modelling
by: Pellegrino, Alessio, et al.
Published: (2025)
by: Pellegrino, Alessio, et al.
Published: (2025)
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use
by: Patel, Hitesh Laxmichand, et al.
Published: (2025)
by: Patel, Hitesh Laxmichand, et al.
Published: (2025)
Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection
by: Selvaganapathy, Sanjeeevan, et al.
Published: (2025)
by: Selvaganapathy, Sanjeeevan, et al.
Published: (2025)
University of Indonesia at SemEval-2025 Task 11: Evaluating State-of-the-Art Encoders for Multi-Label Emotion Detection
by: Hanif, Ikhlasul Akmal, et al.
Published: (2025)
by: Hanif, Ikhlasul Akmal, et al.
Published: (2025)
Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels
by: Rath, Plawan Kumar, et al.
Published: (2026)
by: Rath, Plawan Kumar, et al.
Published: (2026)
Meta-Learning at Scale for Large Language Models via Low-Rank Amortized Bayesian Meta-Learning
by: Zhang, Liyi, et al.
Published: (2025)
by: Zhang, Liyi, et al.
Published: (2025)
LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance
by: Ivanov, Igor
Published: (2025)
by: Ivanov, Igor
Published: (2025)
How LLMs Are Persuaded: A Few Attention Heads, Rerouted
by: Sun, Xiangkun, et al.
Published: (2026)
by: Sun, Xiangkun, et al.
Published: (2026)
PersistBench: When Should Long-Term Memories Be Forgotten by LLMs?
by: Pulipaka, Sidharth, et al.
Published: (2026)
by: Pulipaka, Sidharth, et al.
Published: (2026)
Understanding LLM Evaluator Behavior: A Structured Multi-Evaluator Framework for Merchant Risk Assessment
by: Wang, Liang, et al.
Published: (2026)
by: Wang, Liang, et al.
Published: (2026)
An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs
by: Rai, Daking, et al.
Published: (2024)
by: Rai, Daking, et al.
Published: (2024)
AI Predicts AGI: Leveraging AGI Forecasting and Peer Review to Explore LLMs' Complex Reasoning Capabilities
by: Davide, Fabrizio, et al.
Published: (2024)
by: Davide, Fabrizio, et al.
Published: (2024)
A Performance Evaluation of a Quantized Large Language Model on Various Smartphones
by: Çöplü, Tolga, et al.
Published: (2023)
by: Çöplü, Tolga, et al.
Published: (2023)
Evaluating Steering Techniques using Human Similarity Judgments
by: Studdiford, Zach, et al.
Published: (2025)
by: Studdiford, Zach, et al.
Published: (2025)
Intrinsic Evaluation of RAG Systems for Deep-Logic Questions
by: Hu, Junyi, et al.
Published: (2024)
by: Hu, Junyi, et al.
Published: (2024)
Evaluating LLM Metrics Through Real-World Capabilities
by: Miller, Justin K, et al.
Published: (2025)
by: Miller, Justin K, et al.
Published: (2025)
Annif at SemEval-2025 Task 5: Traditional XMTC augmented by LLMs
by: Suominen, Osma, et al.
Published: (2025)
by: Suominen, Osma, et al.
Published: (2025)
Large Language Model (LLM) Bias Index -- LLMBI
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
BitCal-TTS: Bit-Calibrated Test-Time Scaling for Quantized Reasoning Models
by: Patarlapalli, Sai Babu, et al.
Published: (2026)
by: Patarlapalli, Sai Babu, et al.
Published: (2026)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMs
by: Iqbal, Hasan, et al.
Published: (2024)
by: Iqbal, Hasan, et al.
Published: (2024)
The Unreasonable Effectiveness of Model Merging for Cross-Lingual Transfer in LLMs
by: Bandarkar, Lucas, et al.
Published: (2025)
by: Bandarkar, Lucas, et al.
Published: (2025)
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning
by: Mircea, Andrei, et al.
Published: (2025)
by: Mircea, Andrei, et al.
Published: (2025)
CoDA: Coding LM via Diffusion Adaptation
by: Chen, Haolin, et al.
Published: (2025)
by: Chen, Haolin, et al.
Published: (2025)
Similar Items
-
Allocate Marginal Reviews to Borderline Papers Using LLM Comparative Ranking
by: Epstein, Elliot L., et al.
Published: (2026) -
Do Language Models Mirror Human Confidence? Exploring Psychological Insights to Address Overconfidence in LLMs
by: Xu, Chenjun, et al.
Published: (2025) -
SD-KDE: Score-Debiased Kernel Density Estimation
by: Epstein, Elliot L., et al.
Published: (2025) -
Flash-SD-KDE: Accelerating SD-KDE with Tensor Cores
by: Epstein, Elliot L., et al.
Published: (2026) -
The Metacognitive Probe: Five Behavioural Calibration Diagnostics for LLMs
by: Oliveira, Rafael C. T.
Published: (2026)