Process Supervision of Confidence Margin for Calibrated LLM Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Liaoyaqi, Zuo, Chunsheng, Jurayj, William, Van Durme, Benjamin, Liu, Anqi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Always Tell Me The Odds: Fine-grained Conditional Probability Estimation
by: Wang, Liaoyaqi, et al.
Published: (2025)
by: Wang, Liaoyaqi, et al.
Published: (2025)
Conformal Linguistic Calibration: Trading-off between Factuality and Specificity
by: Jiang, Zhengping, et al.
Published: (2025)
by: Jiang, Zhengping, et al.
Published: (2025)
Language Models and Logic Programs for Trustworthy Tax Reasoning
by: Jurayj, William, et al.
Published: (2025)
by: Jurayj, William, et al.
Published: (2025)
MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
by: Subramani, Nishant, et al.
Published: (2025)
by: Subramani, Nishant, et al.
Published: (2025)
Is That Your Final Answer? Test-Time Scaling Improves Selective Question Answering
by: Jurayj, William, et al.
Published: (2025)
by: Jurayj, William, et al.
Published: (2025)
SEQR: Secure and Efficient QR-based LoRA Routing
by: Fleshman, William, et al.
Published: (2025)
by: Fleshman, William, et al.
Published: (2025)
LoRA-Augmented Generation (LAG) for Knowledge-Intensive Language Tasks
by: Fleshman, William, et al.
Published: (2025)
by: Fleshman, William, et al.
Published: (2025)
RE-Adapt: Reverse Engineered Adaptation of Large Language Models
by: Fleshman, William, et al.
Published: (2024)
by: Fleshman, William, et al.
Published: (2024)
SpectR: Dynamically Composing LM Experts with Spectral Routing
by: Fleshman, William, et al.
Published: (2025)
by: Fleshman, William, et al.
Published: (2025)
RE-AdaptIR: Improving Information Retrieval through Reverse Engineered Adaptation
by: Fleshman, William, et al.
Published: (2024)
by: Fleshman, William, et al.
Published: (2024)
Breaking Symmetry When Training Transformers
by: Zuo, Chunsheng, et al.
Published: (2024)
by: Zuo, Chunsheng, et al.
Published: (2024)
The NLP Task Effectiveness of Long-Range Transformers
by: Qin, Guanghui, et al.
Published: (2022)
by: Qin, Guanghui, et al.
Published: (2022)
Confidence over Time: Confidence Calibration with Temporal Logic for Large Language Model Reasoning
by: Mao, Zhenjiang, et al.
Published: (2026)
by: Mao, Zhenjiang, et al.
Published: (2026)
AdapterSwap: Continuous Training of LLMs with Data Removal and Access-Control Guarantees
by: Fleshman, William, et al.
Published: (2024)
by: Fleshman, William, et al.
Published: (2024)
Many-Tier Instruction Hierarchy in LLM Agents
by: Zhang, Jingyu, et al.
Published: (2026)
by: Zhang, Jingyu, et al.
Published: (2026)
Position Information Emerges in Causal Transformers Without Positional Encodings via Similarity of Nearby Embeddings
by: Zuo, Chunsheng, et al.
Published: (2024)
by: Zuo, Chunsheng, et al.
Published: (2024)
Unified Multimodal Uncertain Inference
by: Zhang, Dengjia, et al.
Published: (2026)
by: Zhang, Dengjia, et al.
Published: (2026)
DeonticBench: A Benchmark for Reasoning over Rules
by: Dou, Guangyao, et al.
Published: (2026)
by: Dou, Guangyao, et al.
Published: (2026)
Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
by: Liu, Haolin, et al.
Published: (2026)
by: Liu, Haolin, et al.
Published: (2026)
Know When You're Wrong: Aligning Confidence with Correctness for LLM Error Detection
by: Xiaohu, Xie, et al.
Published: (2026)
by: Xiaohu, Xie, et al.
Published: (2026)
Weird Generalization is Weirdly Brittle
by: Wanner, Miriam, et al.
Published: (2026)
by: Wanner, Miriam, et al.
Published: (2026)
QA-Calibration of Language Model Confidence Scores
by: Manggala, Putra, et al.
Published: (2024)
by: Manggala, Putra, et al.
Published: (2024)
From the Inside Out: Progressive Distribution Refinement for Confidence Calibration
by: Yang, Xizhong, et al.
Published: (2026)
by: Yang, Xizhong, et al.
Published: (2026)
Do Androids Know They're Only Dreaming of Electric Sheep?
by: CH-Wang, Sky, et al.
Published: (2023)
by: CH-Wang, Sky, et al.
Published: (2023)
Improve Mathematical Reasoning in Language Models by Automated Process Supervision
by: Luo, Liangchen, et al.
Published: (2024)
by: Luo, Liangchen, et al.
Published: (2024)
Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards
by: Ma, Zhengzhao, et al.
Published: (2026)
by: Ma, Zhengzhao, et al.
Published: (2026)
Are LLM Decisions Faithful to Verbal Confidence?
by: Wang, Jiawei, et al.
Published: (2026)
by: Wang, Jiawei, et al.
Published: (2026)
LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error
by: Wang, Boshi, et al.
Published: (2024)
by: Wang, Boshi, et al.
Published: (2024)
Gaps or Hallucinations? Gazing into Machine-Generated Legal Analysis for Fine-grained Text Evaluations
by: Hou, Abe Bohan, et al.
Published: (2024)
by: Hou, Abe Bohan, et al.
Published: (2024)
mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
by: Marone, Marc, et al.
Published: (2025)
by: Marone, Marc, et al.
Published: (2025)
Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting
by: Hu, Michael Y., et al.
Published: (2025)
by: Hu, Michael Y., et al.
Published: (2025)
Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
by: Liu, Runze, et al.
Published: (2025)
by: Liu, Runze, et al.
Published: (2025)
Confidence Calibration in Large Language Model-Based Entity Matching
by: Kamsteeg, Iris, et al.
Published: (2025)
by: Kamsteeg, Iris, et al.
Published: (2025)
Confidence in the Reasoning of Large Language Models
by: Pawitan, Yudi, et al.
Published: (2024)
by: Pawitan, Yudi, et al.
Published: (2024)
Multimodal Clinical Trial Outcome Prediction with Large Language Models
by: Zheng, Wenhao, et al.
Published: (2024)
by: Zheng, Wenhao, et al.
Published: (2024)
IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning
by: Mishra, Aayush, et al.
Published: (2025)
by: Mishra, Aayush, et al.
Published: (2025)
Confidence Geometry Reveals Trace-Level Correctness in Large Language Model Reasoning
by: Liu, Shuo, et al.
Published: (2026)
by: Liu, Shuo, et al.
Published: (2026)
Graph-based Confidence Calibration for Large Language Models
by: Li, Yukun, et al.
Published: (2024)
by: Li, Yukun, et al.
Published: (2024)
RATIONALYST: Mining Implicit Rationales for Process Supervision of Reasoning
by: Jiang, Dongwei, et al.
Published: (2024)
by: Jiang, Dongwei, et al.
Published: (2024)
Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning
by: Ding, Fei, et al.
Published: (2026)
by: Ding, Fei, et al.
Published: (2026)
Similar Items
-
Always Tell Me The Odds: Fine-grained Conditional Probability Estimation
by: Wang, Liaoyaqi, et al.
Published: (2025) -
Conformal Linguistic Calibration: Trading-off between Factuality and Specificity
by: Jiang, Zhengping, et al.
Published: (2025) -
Language Models and Logic Programs for Trustworthy Tax Reasoning
by: Jurayj, William, et al.
Published: (2025) -
MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
by: Subramani, Nishant, et al.
Published: (2025) -
Is That Your Final Answer? Test-Time Scaling Improves Selective Question Answering
by: Jurayj, William, et al.
Published: (2025)