Confidence Estimation for LLMs in Multi-turn Interactions
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Caiqi, Yang, Ruihan, Zhu, Xiaochen, Li, Chengzu, Hu, Tiancheng, Dong, Yijiang River, Yang, Deqing, Collier, Nigel |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations
by: Zhang, Caiqi, et al.
Published: (2025)
by: Zhang, Caiqi, et al.
Published: (2025)
Can LLM be a Personalized Judge?
by: Dong, Yijiang River, et al.
Published: (2024)
by: Dong, Yijiang River, et al.
Published: (2024)
Steer Model beyond Assistant: Controlling System Prompt Strength via Contrastive Decoding
by: Dong, Yijiang River, et al.
Published: (2026)
by: Dong, Yijiang River, et al.
Published: (2026)
Value of Information: A Framework for Human-Agent Communication
by: Dong, Yijiang River, et al.
Published: (2026)
by: Dong, Yijiang River, et al.
Published: (2026)
When Personalization Meets Reality: A Multi-Faceted Analysis of Personalized Preference Learning
by: Dong, Yijiang River, et al.
Published: (2025)
by: Dong, Yijiang River, et al.
Published: (2025)
UNCLE: Benchmarking Uncertainty Expressions in Long-Form Generation
by: Yang, Ruihan, et al.
Published: (2025)
by: Yang, Ruihan, et al.
Published: (2025)
Beyond the Final Layer: Intermediate Representations for Better Multilingual Calibration in Large Language Models
by: Zhou, Ej, et al.
Published: (2025)
by: Zhou, Ej, et al.
Published: (2025)
Atomic Calibration of LLMs in Long-Form Generations
by: Zhang, Caiqi, et al.
Published: (2024)
by: Zhang, Caiqi, et al.
Published: (2024)
LoGU: Long-form Generation with Uncertainty Expressions
by: Yang, Ruihan, et al.
Published: (2024)
by: Yang, Ruihan, et al.
Published: (2024)
All Roads Lead to Rome: Graph-Based Confidence Estimation for Large Language Model Reasoning
by: Zhang, Caiqi, et al.
Published: (2025)
by: Zhang, Caiqi, et al.
Published: (2025)
TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
by: Hui, Zheng, et al.
Published: (2025)
by: Hui, Zheng, et al.
Published: (2025)
Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning
by: Hui, Zheng, et al.
Published: (2025)
by: Hui, Zheng, et al.
Published: (2025)
TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
by: Li, Chengzu, et al.
Published: (2024)
by: Li, Chengzu, et al.
Published: (2024)
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
by: Sivapiromrat, Sanhanat, et al.
Published: (2025)
by: Sivapiromrat, Sanhanat, et al.
Published: (2025)
Conformity in Large Language Models
by: Zhu, Xiaochen, et al.
Published: (2024)
by: Zhu, Xiaochen, et al.
Published: (2024)
LUQ: Long-text Uncertainty Quantification for LLMs
by: Zhang, Caiqi, et al.
Published: (2024)
by: Zhang, Caiqi, et al.
Published: (2024)
iNews: A Multimodal Dataset for Modeling Personalized Affective Responses to News
by: Hu, Tiancheng, et al.
Published: (2025)
by: Hu, Tiancheng, et al.
Published: (2025)
Quantifying the Persona Effect in LLM Simulations
by: Hu, Tiancheng, et al.
Published: (2024)
by: Hu, Tiancheng, et al.
Published: (2024)
Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging
by: Hu, Tiancheng, et al.
Published: (2025)
by: Hu, Tiancheng, et al.
Published: (2025)
Lost in Embeddings: Information Loss in Vision-Language Models
by: Li, Wenyan, et al.
Published: (2025)
by: Li, Wenyan, et al.
Published: (2025)
Evaluating and Enhancing LLMs for Multi-turn Text-to-SQL with Multiple Question Types
by: Guo, Ziming, et al.
Published: (2024)
by: Guo, Ziming, et al.
Published: (2024)
Decoupling the Effect of Chain-of-Thought Reasoning: A Human Label Variation Perspective
by: Chen, Beiduo, et al.
Published: (2026)
by: Chen, Beiduo, et al.
Published: (2026)
Multi-agent AI systems outperform human teams in creativity
by: Hu, Tiancheng, et al.
Published: (2026)
by: Hu, Tiancheng, et al.
Published: (2026)
Generative Language Models Exhibit Social Identity Biases
by: Hu, Tiancheng, et al.
Published: (2023)
by: Hu, Tiancheng, et al.
Published: (2023)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
by: Hu, Tiancheng, et al.
Published: (2025)
by: Hu, Tiancheng, et al.
Published: (2025)
UNDIAL: Self-Distillation with Adjusted Logits for Robust Unlearning in Large Language Models
by: Dong, Yijiang River, et al.
Published: (2024)
by: Dong, Yijiang River, et al.
Published: (2024)
Confidence Calibration and Rationalization for LLMs via Multi-Agent Deliberation
by: Yang, Ruixin, et al.
Published: (2024)
by: Yang, Ruixin, et al.
Published: (2024)
Visual Planning: Let's Think Only with Images
by: Xu, Yi, et al.
Published: (2025)
by: Xu, Yi, et al.
Published: (2025)
PiVe: Prompting with Iterative Verification Improving Graph-based Generative Capability of LLMs
by: Han, Jiuzhou, et al.
Published: (2023)
by: Han, Jiuzhou, et al.
Published: (2023)
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
by: Balkır, Esma, et al.
Published: (2026)
by: Balkır, Esma, et al.
Published: (2026)
Attention Instruction: Amplifying Attention in the Middle via Prompting
by: Zhang, Meiru, et al.
Published: (2024)
by: Zhang, Meiru, et al.
Published: (2024)
A Survey on Prompt Tuning
by: Li, Zongqian, et al.
Published: (2025)
by: Li, Zongqian, et al.
Published: (2025)
PT-MoE: An Efficient Finetuning Framework for Integrating Mixture-of-Experts into Prompt Tuning
by: Li, Zongqian, et al.
Published: (2025)
by: Li, Zongqian, et al.
Published: (2025)
500xCompressor: Generalized Prompt Compression for Large Language Models
by: Li, Zongqian, et al.
Published: (2024)
by: Li, Zongqian, et al.
Published: (2024)
Facilitating Multi-turn Function Calling for LLMs via Compositional Instruction Tuning
by: Chen, Mingyang, et al.
Published: (2024)
by: Chen, Mingyang, et al.
Published: (2024)
FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMs
by: Fan, Zhiting, et al.
Published: (2024)
by: Fan, Zhiting, et al.
Published: (2024)
On Verbalized Confidence Scores for LLMs
by: Yang, Daniel, et al.
Published: (2024)
by: Yang, Daniel, et al.
Published: (2024)
Time to Revist Exact Match
by: Abbood, Auss, et al.
Published: (2025)
by: Abbood, Auss, et al.
Published: (2025)
Multi-lingual Multi-turn Automated Red Teaming for LLMs
by: Singhania, Abhishek, et al.
Published: (2025)
by: Singhania, Abhishek, et al.
Published: (2025)
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback
by: Wang, Xingyao, et al.
Published: (2023)
by: Wang, Xingyao, et al.
Published: (2023)
Similar Items
-
LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations
by: Zhang, Caiqi, et al.
Published: (2025) -
Can LLM be a Personalized Judge?
by: Dong, Yijiang River, et al.
Published: (2024) -
Steer Model beyond Assistant: Controlling System Prompt Strength via Contrastive Decoding
by: Dong, Yijiang River, et al.
Published: (2026) -
Value of Information: A Framework for Human-Agent Communication
by: Dong, Yijiang River, et al.
Published: (2026) -
When Personalization Meets Reality: A Multi-Faceted Analysis of Personalized Preference Learning
by: Dong, Yijiang River, et al.
Published: (2025)