Confidence Calibration and Rationalization for LLMs via Multi-Agent Deliberation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Ruixin, Rajagopal, Dheeraj, Hayati, Shirley Anugrah, Hu, Bin, Kang, Dongyeop |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Far Can We Extract Diverse Perspectives from Large Language Models?
von: Hayati, Shirley Anugrah, et al.
Veröffentlicht: (2023)
von: Hayati, Shirley Anugrah, et al.
Veröffentlicht: (2023)
Effects of Varying LLM Access on Essay Writing Behavior
von: Christenson, Julia, et al.
Veröffentlicht: (2026)
von: Christenson, Julia, et al.
Veröffentlicht: (2026)
Are LLM Agents Behaviorally Coherent? Latent Profiles for Social Simulation
von: Mooney, James, et al.
Veröffentlicht: (2025)
von: Mooney, James, et al.
Veröffentlicht: (2025)
Chain-of-Instructions: Compositional Instruction Tuning on Large Language Models
von: Hayati, Shirley Anugrah, et al.
Veröffentlicht: (2024)
von: Hayati, Shirley Anugrah, et al.
Veröffentlicht: (2024)
Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL
von: Zhai, Skylar, et al.
Veröffentlicht: (2026)
von: Zhai, Skylar, et al.
Veröffentlicht: (2026)
Learning to Negotiate: Multi-Agent Deliberation for Collective Value Alignment in LLMs
von: Anantaprayoon, Panatchakorn, et al.
Veröffentlicht: (2026)
von: Anantaprayoon, Panatchakorn, et al.
Veröffentlicht: (2026)
Dynamic Multi-Reward Weighting for Multi-Style Controllable Generation
von: de Langis, Karin, et al.
Veröffentlicht: (2024)
von: de Langis, Karin, et al.
Veröffentlicht: (2024)
Double-Calibration: Towards Reliable LLMs via Calibrating Knowledge and Reasoning Confidence
von: Lu, Yuyin, et al.
Veröffentlicht: (2026)
von: Lu, Yuyin, et al.
Veröffentlicht: (2026)
Confidence Estimation for LLMs in Multi-turn Interactions
von: Zhang, Caiqi, et al.
Veröffentlicht: (2026)
von: Zhang, Caiqi, et al.
Veröffentlicht: (2026)
Under the Surface: Tracking the Artifactuality of LLM-Generated Data
von: Das, Debarati, et al.
Veröffentlicht: (2024)
von: Das, Debarati, et al.
Veröffentlicht: (2024)
Enhancing Language Model Rationality with Bi-Directional Deliberation Reasoning
von: Zhang, Yadong, et al.
Veröffentlicht: (2024)
von: Zhang, Yadong, et al.
Veröffentlicht: (2024)
How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMs
von: de Langis, Karin, et al.
Veröffentlicht: (2025)
von: de Langis, Karin, et al.
Veröffentlicht: (2025)
Strong Memory, Weak Control: An Empirical Study of Executive Functioning in LLMs
von: de Langis, Karin, et al.
Veröffentlicht: (2025)
von: de Langis, Karin, et al.
Veröffentlicht: (2025)
SelectLLM: Can LLMs Select Important Instructions to Annotate?
von: Parkar, Ritik Sachin, et al.
Veröffentlicht: (2024)
von: Parkar, Ritik Sachin, et al.
Veröffentlicht: (2024)
NAACL: Noise-AwAre Verbal Confidence Calibration for Robust LLMs in RAG Systems
von: Liu, Jiayu, et al.
Veröffentlicht: (2026)
von: Liu, Jiayu, et al.
Veröffentlicht: (2026)
II-MMR: Identifying and Improving Multi-modal Multi-hop Reasoning in Visual Question Answering
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
von: Zhang, Zhiwei, et al.
Veröffentlicht: (2025)
von: Zhang, Zhiwei, et al.
Veröffentlicht: (2025)
Achieving Unanimous Consensus Through Multi-Agent Deliberation
von: Pokharel, Apurba, et al.
Veröffentlicht: (2025)
von: Pokharel, Apurba, et al.
Veröffentlicht: (2025)
BaseCal: Unsupervised Confidence Calibration via Base Model Signals
von: Tan, Hexiang, et al.
Veröffentlicht: (2026)
von: Tan, Hexiang, et al.
Veröffentlicht: (2026)
Scalable Influence and Fact Tracing for Large Language Model Pretraining
von: Chang, Tyler A., et al.
Veröffentlicht: (2024)
von: Chang, Tyler A., et al.
Veröffentlicht: (2024)
CascadeDebate: Multi-Agent Deliberation for Cost-Aware LLM Cascades
von: Chang, Raeyoung, et al.
Veröffentlicht: (2026)
von: Chang, Raeyoung, et al.
Veröffentlicht: (2026)
On Verbalized Confidence Scores for LLMs
von: Yang, Daniel, et al.
Veröffentlicht: (2024)
von: Yang, Daniel, et al.
Veröffentlicht: (2024)
Agentic Confidence Calibration
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2026)
Enhancing Multi-Agent Debate System Performance via Confidence Expression
von: Lin, Zijie, et al.
Veröffentlicht: (2025)
von: Lin, Zijie, et al.
Veröffentlicht: (2025)
Steering off Course: Reliability Challenges in Steering Language Models
von: Da Silva, Patrick Queiroz, et al.
Veröffentlicht: (2025)
von: Da Silva, Patrick Queiroz, et al.
Veröffentlicht: (2025)
Identifying Influential N-grams in Confidence Calibration via Regression Analysis
von: Ozaki, Shintaro, et al.
Veröffentlicht: (2026)
von: Ozaki, Shintaro, et al.
Veröffentlicht: (2026)
SAND: Boosting LLM Agents with Self-Taught Action Deliberation
von: Xia, Yu, et al.
Veröffentlicht: (2025)
von: Xia, Yu, et al.
Veröffentlicht: (2025)
Calibrated Confidence Expression for Radiology Report Generation
von: Bani-Harouni, David, et al.
Veröffentlicht: (2026)
von: Bani-Harouni, David, et al.
Veröffentlicht: (2026)
Judging with Confidence: Calibrating Autoraters to Preference Distributions
von: Li, Zhuohang, et al.
Veröffentlicht: (2025)
von: Li, Zhuohang, et al.
Veröffentlicht: (2025)
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
von: Xiong, Miao, et al.
Veröffentlicht: (2023)
von: Xiong, Miao, et al.
Veröffentlicht: (2023)
Mary, the Cheeseburger-Eating Vegetarian: Do LLMs Recognize Incoherence in Narratives?
von: de Langis, Karin, et al.
Veröffentlicht: (2025)
von: de Langis, Karin, et al.
Veröffentlicht: (2025)
Teaching LLMs to Abstain via Fine-Grained Semantic Confidence Reward
von: An, Hao, et al.
Veröffentlicht: (2025)
von: An, Hao, et al.
Veröffentlicht: (2025)
Found in Conversation: LLMs Teach Themselves to Close the Multi-Turn Gap
von: Chen, Tianlang, et al.
Veröffentlicht: (2026)
von: Chen, Tianlang, et al.
Veröffentlicht: (2026)
FINEREASON: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle Solving
von: Chen, Guizhen, et al.
Veröffentlicht: (2025)
von: Chen, Guizhen, et al.
Veröffentlicht: (2025)
Reasoning over Boundaries: Enhancing Specification Alignment via Test-time Deliberation
von: Zhang, Haoran, et al.
Veröffentlicht: (2025)
von: Zhang, Haoran, et al.
Veröffentlicht: (2025)
It Helps to Take a Second Opinion: Teaching Smaller LLMs to Deliberate Mutually via Selective Rationale Optimisation
von: Patnaik, Sohan, et al.
Veröffentlicht: (2025)
von: Patnaik, Sohan, et al.
Veröffentlicht: (2025)
MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
von: Subramani, Nishant, et al.
Veröffentlicht: (2025)
von: Subramani, Nishant, et al.
Veröffentlicht: (2025)
The Amazing Agent Race: Strong Tool Users, Weak Navigators
von: Kim, Zae Myung, et al.
Veröffentlicht: (2026)
von: Kim, Zae Myung, et al.
Veröffentlicht: (2026)
Enhancing Language Model Factuality via Activation-Based Confidence Calibration and Guided Decoding
von: Liu, Xin, et al.
Veröffentlicht: (2024)
von: Liu, Xin, et al.
Veröffentlicht: (2024)
Calibrating the Confidence of Large Language Models by Eliciting Fidelity
von: Zhang, Mozhi, et al.
Veröffentlicht: (2024)
von: Zhang, Mozhi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
How Far Can We Extract Diverse Perspectives from Large Language Models?
von: Hayati, Shirley Anugrah, et al.
Veröffentlicht: (2023) -
Effects of Varying LLM Access on Essay Writing Behavior
von: Christenson, Julia, et al.
Veröffentlicht: (2026) -
Are LLM Agents Behaviorally Coherent? Latent Profiles for Social Simulation
von: Mooney, James, et al.
Veröffentlicht: (2025) -
Chain-of-Instructions: Compositional Instruction Tuning on Large Language Models
von: Hayati, Shirley Anugrah, et al.
Veröffentlicht: (2024) -
Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL
von: Zhai, Skylar, et al.
Veröffentlicht: (2026)