Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A
Fuente:
arXiv
Saved in:
| Main Authors: | Plaut, Benjamin, Khanh, Nguyen X., Trinh, Tu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Getting By Goal Misgeneralization With a Little Help From a Mentor
by: Trinh, Tu, et al.
Published: (2024)
by: Trinh, Tu, et al.
Published: (2024)
False Fixed Points: Kantian Feedback, Stable Miscalibration, and Representational Compression in LLMs
by: Okutomi, Akira
Published: (2025)
by: Okutomi, Akira
Published: (2025)
Large Language Models are Miscalibrated In-Context Learners
by: Li, Chengzu, et al.
Published: (2023)
by: Li, Chengzu, et al.
Published: (2023)
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction
by: Lee, Yooseop, et al.
Published: (2025)
by: Lee, Yooseop, et al.
Published: (2025)
UnibucLLM: Harnessing LLMs for Automated Prediction of Item Difficulty and Response Time for Multiple-Choice Questions
by: Rogoz, Ana-Cristina, et al.
Published: (2024)
by: Rogoz, Ana-Cristina, et al.
Published: (2024)
YRC-Bench: A Benchmark for Learning to Coordinate with Experts
by: Danesh, Mohamad H., et al.
Published: (2025)
by: Danesh, Mohamad H., et al.
Published: (2025)
Safety Training Persists Through Helpfulness Optimization in LLM Agents
by: Plaut, Benjamin
Published: (2026)
by: Plaut, Benjamin
Published: (2026)
Frontier LLMs Still Struggle with Simple Reasoning Tasks
by: Malek, Alan, et al.
Published: (2025)
by: Malek, Alan, et al.
Published: (2025)
On Overcoming Miscalibrated Conversational Priors in LLM-based Chatbots
by: Herlihy, Christine, et al.
Published: (2024)
by: Herlihy, Christine, et al.
Published: (2024)
Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting
by: To, Bang Trinh Tran, et al.
Published: (2025)
by: To, Bang Trinh Tran, et al.
Published: (2025)
Biomedical Entity Linking as Multiple Choice Question Answering
by: Lin, Zhenxi, et al.
Published: (2024)
by: Lin, Zhenxi, et al.
Published: (2024)
Option-ID Based Elimination For Multiple Choice Questions
by: Zhu, Zhenhao, et al.
Published: (2025)
by: Zhu, Zhenhao, et al.
Published: (2025)
Exploring Design Choices for Building Language-Specific LLMs
by: Tejaswi, Atula, et al.
Published: (2024)
by: Tejaswi, Atula, et al.
Published: (2024)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
by: Sun, Yiyou, et al.
Published: (2025)
by: Sun, Yiyou, et al.
Published: (2025)
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
by: Chandak, Nikhil, et al.
Published: (2025)
by: Chandak, Nikhil, et al.
Published: (2025)
Multiple Choice Learning of Low-Rank Adapters for Language Modeling
by: Letzelter, Victor, et al.
Published: (2025)
by: Letzelter, Victor, et al.
Published: (2025)
A Study on Large Language Models' Limitations in Multiple-Choice Question Answering
by: Khatun, Aisha, et al.
Published: (2024)
by: Khatun, Aisha, et al.
Published: (2024)
Pattern Recognition or Medical Knowledge? The Problem with Multiple-Choice Questions in Medicine
by: Griot, Maxime, et al.
Published: (2024)
by: Griot, Maxime, et al.
Published: (2024)
Federated Document Visual Question Answering: A Pilot Study
by: Nguyen, Khanh, et al.
Published: (2024)
by: Nguyen, Khanh, et al.
Published: (2024)
Learning When Not to Learn: Risk-Sensitive Abstention in Bandits with Unbounded Rewards
by: Liaw, Sarah, et al.
Published: (2025)
by: Liaw, Sarah, et al.
Published: (2025)
Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions
by: Li, Ruizhe, et al.
Published: (2024)
by: Li, Ruizhe, et al.
Published: (2024)
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
by: Yang, Zhihe, et al.
Published: (2025)
by: Yang, Zhihe, et al.
Published: (2025)
Language-Guided World Models: A Model-Based Approach to AI Control
by: Zhang, Alex, et al.
Published: (2024)
by: Zhang, Alex, et al.
Published: (2024)
A Course Correction in Steerability Evaluation: Revealing Miscalibration and Side Effects in LLMs
by: Chang, Trenton, et al.
Published: (2025)
by: Chang, Trenton, et al.
Published: (2025)
Evaluating the Performance and Robustness of LLMs in Materials Science Q&A and Property Predictions
by: Wang, Hongchen, et al.
Published: (2024)
by: Wang, Hongchen, et al.
Published: (2024)
Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs
by: Goel, Raghavv, et al.
Published: (2024)
by: Goel, Raghavv, et al.
Published: (2024)
Always Tell Me The Odds: Fine-grained Conditional Probability Estimation
by: Wang, Liaoyaqi, et al.
Published: (2025)
by: Wang, Liaoyaqi, et al.
Published: (2025)
Tree-OPO: Off-policy Monte Carlo Tree-Guided Advantage Optimization for Multistep Reasoning
by: Huang, Bingning, et al.
Published: (2025)
by: Huang, Bingning, et al.
Published: (2025)
Generating Multiple-Choice Knowledge Questions with Interpretable Difficulty Estimation using Knowledge Graphs and Large Language Models
by: Şakiroğlu, Mehmet Can, et al.
Published: (2026)
by: Şakiroğlu, Mehmet Can, et al.
Published: (2026)
LoRMA: Low-Rank Multiplicative Adaptation for LLMs
by: Bihany, Harsh, et al.
Published: (2025)
by: Bihany, Harsh, et al.
Published: (2025)
Learning to Correct for QA Reasoning with Black-box LLMs
by: Kim, Jaehyung, et al.
Published: (2024)
by: Kim, Jaehyung, et al.
Published: (2024)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
by: Lin, Zicheng, et al.
Published: (2024)
by: Lin, Zicheng, et al.
Published: (2024)
Still "Talking About Large Language Models": Some Clarifications
by: Shanahan, Murray
Published: (2024)
by: Shanahan, Murray
Published: (2024)
Small Models Are (Still) Effective Cross-Domain Argument Extractors
by: Gantt, William, et al.
Published: (2024)
by: Gantt, William, et al.
Published: (2024)
Tagengo: A Multilingual Chat Dataset
by: Devine, Peter
Published: (2024)
by: Devine, Peter
Published: (2024)
Efficient Joint Prediction of Multiple Future Tokens
by: Ahn, Kwangjun, et al.
Published: (2025)
by: Ahn, Kwangjun, et al.
Published: (2025)
Scaling Laws for Predicting Downstream Performance in LLMs
by: Chen, Yangyi, et al.
Published: (2024)
by: Chen, Yangyi, et al.
Published: (2024)
Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting
by: Liu, Chi, et al.
Published: (2026)
by: Liu, Chi, et al.
Published: (2026)
Teaching LLMs for Step-Level Automatic Math Correction via Reinforcement Learning
by: Li, Junsong, et al.
Published: (2025)
by: Li, Junsong, et al.
Published: (2025)
Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability
by: Lee, Yu-Ting, et al.
Published: (2025)
by: Lee, Yu-Ting, et al.
Published: (2025)
Similar Items
-
Getting By Goal Misgeneralization With a Little Help From a Mentor
by: Trinh, Tu, et al.
Published: (2024) -
False Fixed Points: Kantian Feedback, Stable Miscalibration, and Representational Compression in LLMs
by: Okutomi, Akira
Published: (2025) -
Large Language Models are Miscalibrated In-Context Learners
by: Li, Chengzu, et al.
Published: (2023) -
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction
by: Lee, Yooseop, et al.
Published: (2025) -
UnibucLLM: Harnessing LLMs for Automated Prediction of Item Difficulty and Response Time for Multiple-Choice Questions
by: Rogoz, Ana-Cristina, et al.
Published: (2024)