Confidence Should Be Calibrated More Than One Turn Deep
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Zhaohan, Li, Chengzhengxu, Liu, Xiaoming, Shen, Chao, Liu, Ziquan, Patras, Ioannis |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models
von: Zhang, Zhaohan, et al.
Veröffentlicht: (2025)
von: Zhang, Zhaohan, et al.
Veröffentlicht: (2025)
Get Confused Cautiously: Textual Sequence Memorization Erasure with Selective Entropy Maximization
von: Zhang, Zhaohan, et al.
Veröffentlicht: (2024)
von: Zhang, Zhaohan, et al.
Veröffentlicht: (2024)
Iron Sharpens Iron: Defending Against Attacks in Machine-Generated Text Detection with Adversarial Training
von: Li, Yuanfan, et al.
Veröffentlicht: (2025)
von: Li, Yuanfan, et al.
Veröffentlicht: (2025)
StablePT: Towards Stable Prompting for Few-shot Learning via Input Separation
von: Liu, Xiaoming, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoming, et al.
Veröffentlicht: (2024)
Concentrate Attention: Towards Domain-Generalizable Prompt Optimization for Language Models
von: Li, Chengzhengxu, et al.
Veröffentlicht: (2024)
von: Li, Chengzhengxu, et al.
Veröffentlicht: (2024)
Does DetectGPT Fully Utilize Perturbation? Bridging Selective Perturbation to Fine-tuned Contrastive Learning Detector would be Better
von: Liu, Shengchao, et al.
Veröffentlicht: (2024)
von: Liu, Shengchao, et al.
Veröffentlicht: (2024)
Upfront Chain-of-Thought: A Cooperative Framework for Chain-of-Thought Compression
von: Li, Chengzhengxu, et al.
Veröffentlicht: (2025)
von: Li, Chengzhengxu, et al.
Veröffentlicht: (2025)
MGTEVAL: An Interactive Platform for Systemtic Evaluation of Machine-Generated Text Detectors
von: Li, Yuanfan, et al.
Veröffentlicht: (2026)
von: Li, Yuanfan, et al.
Veröffentlicht: (2026)
MGT-Prism: Enhancing Domain Generalization for Machine-Generated Text Detection via Spectral Alignment
von: Liu, Shengchao, et al.
Veröffentlicht: (2025)
von: Liu, Shengchao, et al.
Veröffentlicht: (2025)
Safety Alignment Should Be Made More Than Just A Few Attention Heads
von: Huang, Chao, et al.
Veröffentlicht: (2025)
von: Huang, Chao, et al.
Veröffentlicht: (2025)
Dialogue for Prompting: a Policy-Gradient-Based Discrete Prompt Generation for Few-shot Learning
von: Li, Chengzhengxu, et al.
Veröffentlicht: (2023)
von: Li, Chengzhengxu, et al.
Veröffentlicht: (2023)
DEER: Disentangled Mixture of Experts with Instance-Adaptive Routing for Generalizable Machine-Generated Text Detection
von: Ma, Guoxin, et al.
Veröffentlicht: (2025)
von: Ma, Guoxin, et al.
Veröffentlicht: (2025)
Should LLM Safety Be More Than Refusing Harmful Instructions?
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)
Triplet-Structured Knowledge Integration for Multi-Turn Medical Reasoning
von: Meng, Zhaohan, et al.
Veröffentlicht: (2025)
von: Meng, Zhaohan, et al.
Veröffentlicht: (2025)
Long Is More Important Than Difficult for Training Reasoning Models
von: Shen, Si, et al.
Veröffentlicht: (2025)
von: Shen, Si, et al.
Veröffentlicht: (2025)
From Calibration to Collaboration: LLM Uncertainty Quantification Should Be More Human-Centered
von: Devic, Siddartha, et al.
Veröffentlicht: (2025)
von: Devic, Siddartha, et al.
Veröffentlicht: (2025)
Your Large Language Model is Secretly a Fairness Proponent and You Should Prompt it Like One
von: Li, Tianlin, et al.
Veröffentlicht: (2024)
von: Li, Tianlin, et al.
Veröffentlicht: (2024)
Many LLMs Are More Utilitarian Than One
von: Keshmirian, Anita, et al.
Veröffentlicht: (2025)
von: Keshmirian, Anita, et al.
Veröffentlicht: (2025)
Agentic Confidence Calibration
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2026)
Breaking Language Barriers or Reinforcing Bias? A Study of Gender and Racial Disparities in Multilingual Contrastive Vision Language Models
von: Sahili, Zahraa Al, et al.
Veröffentlicht: (2025)
von: Sahili, Zahraa Al, et al.
Veröffentlicht: (2025)
More Than One Teacher: Adaptive Multi-Guidance Policy Optimization for Diverse Exploration
von: Yuan, Xiaoyang, et al.
Veröffentlicht: (2025)
von: Yuan, Xiaoyang, et al.
Veröffentlicht: (2025)
Fact-Level Confidence Calibration and Self-Correction
von: Yuan, Yige, et al.
Veröffentlicht: (2024)
von: Yuan, Yige, et al.
Veröffentlicht: (2024)
MoreHopQA: More Than Multi-hop Reasoning
von: Schnitzler, Julian, et al.
Veröffentlicht: (2024)
von: Schnitzler, Julian, et al.
Veröffentlicht: (2024)
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
Attention Reveals More Than Tokens: Training-Free Long-Context Reasoning with Attention-guided Retrieval
von: Zhang, Yuwei, et al.
Veröffentlicht: (2025)
von: Zhang, Yuwei, et al.
Veröffentlicht: (2025)
Judging with Confidence: Calibrating Autoraters to Preference Distributions
von: Li, Zhuohang, et al.
Veröffentlicht: (2025)
von: Li, Zhuohang, et al.
Veröffentlicht: (2025)
AIM-Fair: Advancing Algorithmic Fairness via Selectively Fine-Tuning Biased Models with Contextual Synthetic Data
von: Zhao, Zengqun, et al.
Veröffentlicht: (2025)
von: Zhao, Zengqun, et al.
Veröffentlicht: (2025)
Rethinking Layer Redundancy: Calibration Matters More Than Search in LLM Depth Pruning
von: Kim, Minkyu, et al.
Veröffentlicht: (2026)
von: Kim, Minkyu, et al.
Veröffentlicht: (2026)
Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations
von: Xia, Yuxi, et al.
Veröffentlicht: (2026)
von: Xia, Yuxi, et al.
Veröffentlicht: (2026)
Self-Anchoring Calibration Drift in Large Language Models: How Multi-Turn Conversations Reshape Model Confidence
von: Harshavardhan
Veröffentlicht: (2026)
von: Harshavardhan
Veröffentlicht: (2026)
CritiCal: Can Critique Help LLM Uncertainty or Confidence Calibration?
von: Zong, Qing, et al.
Veröffentlicht: (2025)
von: Zong, Qing, et al.
Veröffentlicht: (2025)
Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey
von: Liu, Xiaoou, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoou, et al.
Veröffentlicht: (2025)
Process Supervision of Confidence Margin for Calibrated LLM Reasoning
von: Wang, Liaoyaqi, et al.
Veröffentlicht: (2026)
von: Wang, Liaoyaqi, et al.
Veröffentlicht: (2026)
More Than Catastrophic Forgetting: Integrating General Capabilities For Domain-Specific LLMs
von: Liu, Chengyuan, et al.
Veröffentlicht: (2024)
von: Liu, Chengyuan, et al.
Veröffentlicht: (2024)
Graph-based Confidence Calibration for Large Language Models
von: Li, Yukun, et al.
Veröffentlicht: (2024)
von: Li, Yukun, et al.
Veröffentlicht: (2024)
Confidence-Calibrated Small-Large Language Model Collaboration for Cost-Efficient Reasoning
von: Zhang, Chuang, et al.
Veröffentlicht: (2026)
von: Zhang, Chuang, et al.
Veröffentlicht: (2026)
Enhancing Language Model Factuality via Activation-Based Confidence Calibration and Guided Decoding
von: Liu, Xin, et al.
Veröffentlicht: (2024)
von: Liu, Xin, et al.
Veröffentlicht: (2024)
Calibrating the Confidence of Large Language Models by Eliciting Fidelity
von: Zhang, Mozhi, et al.
Veröffentlicht: (2024)
von: Zhang, Mozhi, et al.
Veröffentlicht: (2024)
Trivial Vocabulary Bans Improve LLM Reasoning More Than Deep Linguistic Constraints
von: Jehu-Appiah, Rodney
Veröffentlicht: (2026)
von: Jehu-Appiah, Rodney
Veröffentlicht: (2026)
Calibrated Confidence Expression for Radiology Report Generation
von: Bani-Harouni, David, et al.
Veröffentlicht: (2026)
von: Bani-Harouni, David, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models
von: Zhang, Zhaohan, et al.
Veröffentlicht: (2025) -
Get Confused Cautiously: Textual Sequence Memorization Erasure with Selective Entropy Maximization
von: Zhang, Zhaohan, et al.
Veröffentlicht: (2024) -
Iron Sharpens Iron: Defending Against Attacks in Machine-Generated Text Detection with Adversarial Training
von: Li, Yuanfan, et al.
Veröffentlicht: (2025) -
StablePT: Towards Stable Prompting for Few-shot Learning via Input Separation
von: Liu, Xiaoming, et al.
Veröffentlicht: (2024) -
Concentrate Attention: Towards Domain-Generalizable Prompt Optimization for Language Models
von: Li, Chengzhengxu, et al.
Veröffentlicht: (2024)