LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Caiqi, Zhu, Xiaochen, Li, Chengzu, Collier, Nigel, Vlachos, Andreas |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Conformity in Large Language Models
por: Zhu, Xiaochen, et al.
Publicado: (2024)
por: Zhu, Xiaochen, et al.
Publicado: (2024)
Atomic Calibration of LLMs in Long-Form Generations
por: Zhang, Caiqi, et al.
Publicado: (2024)
por: Zhang, Caiqi, et al.
Publicado: (2024)
All Roads Lead to Rome: Graph-Based Confidence Estimation for Large Language Model Reasoning
por: Zhang, Caiqi, et al.
Publicado: (2025)
por: Zhang, Caiqi, et al.
Publicado: (2025)
LoGU: Long-form Generation with Uncertainty Expressions
por: Yang, Ruihan, et al.
Publicado: (2024)
por: Yang, Ruihan, et al.
Publicado: (2024)
Segment-Level Diffusion: A Framework for Controllable Long-Form Generation with Diffusion Language Models
por: Zhu, Xiaochen, et al.
Publicado: (2024)
por: Zhu, Xiaochen, et al.
Publicado: (2024)
Confidence Estimation for LLMs in Multi-turn Interactions
por: Zhang, Caiqi, et al.
Publicado: (2026)
por: Zhang, Caiqi, et al.
Publicado: (2026)
Beyond the Final Layer: Intermediate Representations for Better Multilingual Calibration in Large Language Models
por: Zhou, Ej, et al.
Publicado: (2025)
por: Zhou, Ej, et al.
Publicado: (2025)
Calibrating Verbalized Confidence with Self-Generated Distractors
por: Wang, Victor, et al.
Publicado: (2025)
por: Wang, Victor, et al.
Publicado: (2025)
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
por: Balkır, Esma, et al.
Publicado: (2026)
por: Balkır, Esma, et al.
Publicado: (2026)
UNCLE: Benchmarking Uncertainty Expressions in Long-Form Generation
por: Yang, Ruihan, et al.
Publicado: (2025)
por: Yang, Ruihan, et al.
Publicado: (2025)
TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
por: Li, Chengzu, et al.
Publicado: (2024)
por: Li, Chengzu, et al.
Publicado: (2024)
LUQ: Long-text Uncertainty Quantification for LLMs
por: Zhang, Caiqi, et al.
Publicado: (2024)
por: Zhang, Caiqi, et al.
Publicado: (2024)
ConfTuner: Training Large Language Models to Express Their Confidence Verbally
por: Li, Yibo, et al.
Publicado: (2025)
por: Li, Yibo, et al.
Publicado: (2025)
Visual Planning: Let's Think Only with Images
por: Xu, Yi, et al.
Publicado: (2025)
por: Xu, Yi, et al.
Publicado: (2025)
How do LLMs Compute Verbal Confidence
por: Kumaran, Dharshan, et al.
Publicado: (2026)
por: Kumaran, Dharshan, et al.
Publicado: (2026)
Collaborative Evaluation of Deepfake Text with Deliberation-Enhancing Dialogue Systems
por: Lee, Jooyoung, et al.
Publicado: (2025)
por: Lee, Jooyoung, et al.
Publicado: (2025)
Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation
por: Yang, Zixuan, et al.
Publicado: (2026)
por: Yang, Zixuan, et al.
Publicado: (2026)
Reasoning Models Better Express Their Confidence
por: Yoon, Dongkeun, et al.
Publicado: (2025)
por: Yoon, Dongkeun, et al.
Publicado: (2025)
ReasonGraph: Visualisation of Reasoning Paths
por: Li, Zongqian, et al.
Publicado: (2025)
por: Li, Zongqian, et al.
Publicado: (2025)
MeVe: A Modular System for Memory Verification and Effective Context Control in Language Models
por: Ottem, Andreas
Publicado: (2025)
por: Ottem, Andreas
Publicado: (2025)
Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging
por: Hu, Tiancheng, et al.
Publicado: (2025)
por: Hu, Tiancheng, et al.
Publicado: (2025)
Verbal Process Supervision Elicits Better Coding Agents
por: Chen, Hao-Yuan, et al.
Publicado: (2025)
por: Chen, Hao-Yuan, et al.
Publicado: (2025)
Verbalized Confidence Triggers Self-Verification: Emergent Behavior Without Explicit Reasoning Supervision
por: Jang, Chaeyun, et al.
Publicado: (2025)
por: Jang, Chaeyun, et al.
Publicado: (2025)
When Personalization Meets Reality: A Multi-Faceted Analysis of Personalized Preference Learning
por: Dong, Yijiang River, et al.
Publicado: (2025)
por: Dong, Yijiang River, et al.
Publicado: (2025)
Do We Need Language-Specific Fact-Checking Models? The Case of Chinese
por: Zhang, Caiqi, et al.
Publicado: (2024)
por: Zhang, Caiqi, et al.
Publicado: (2024)
PaVeRL-SQL: Text-to-SQL via Partial-Match Rewards and Verbal Reinforcement Learning
por: Hao, Heng, et al.
Publicado: (2025)
por: Hao, Heng, et al.
Publicado: (2025)
Knowledge-Level Consistency Reinforcement Learning: Dual-Fact Alignment for Long-Form Factuality
por: Li, Junliang, et al.
Publicado: (2025)
por: Li, Junliang, et al.
Publicado: (2025)
Empathy-R1: A Chain-of-Empathy and Reinforcement Learning Framework for Long-Form Mental Health Support
por: Yao, Xianrong, et al.
Publicado: (2025)
por: Yao, Xianrong, et al.
Publicado: (2025)
Improving Word Translation via Two-Stage Contrastive Learning
por: Li, Yaoyiran, et al.
Publicado: (2022)
por: Li, Yaoyiran, et al.
Publicado: (2022)
Linguistic Calibration of Long-Form Generations
por: Band, Neil, et al.
Publicado: (2024)
por: Band, Neil, et al.
Publicado: (2024)
Semantic Map-based Generation of Navigation Instructions
por: Li, Chengzu, et al.
Publicado: (2024)
por: Li, Chengzu, et al.
Publicado: (2024)
PiVe: Prompting with Iterative Verification Improving Graph-based Generative Capability of LLMs
por: Han, Jiuzhou, et al.
Publicado: (2023)
por: Han, Jiuzhou, et al.
Publicado: (2023)
RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
por: Wu, Mian, et al.
Publicado: (2025)
por: Wu, Mian, et al.
Publicado: (2025)
Integrating Planning into Single-Turn Long-Form Text Generation
por: Liang, Yi, et al.
Publicado: (2024)
por: Liang, Yi, et al.
Publicado: (2024)
An LLM Feature-based Framework for Dialogue Constructiveness Assessment
por: Zhou, Lexin, et al.
Publicado: (2024)
por: Zhou, Lexin, et al.
Publicado: (2024)
Verbal Confidence Saturation in 3-9B Open-Weight Instruction-Tuned LLMs: A Pre-Registered Psychometric Validity Screen
por: Cacioli, Jon-Paul
Publicado: (2026)
por: Cacioli, Jon-Paul
Publicado: (2026)
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
por: Sivapiromrat, Sanhanat, et al.
Publicado: (2025)
por: Sivapiromrat, Sanhanat, et al.
Publicado: (2025)
The Role of Ambiguity in Error Prediction via Uncertainty Quantification
por: Staliūnaitė, Ieva Raminta, et al.
Publicado: (2026)
por: Staliūnaitė, Ieva Raminta, et al.
Publicado: (2026)
Reasoning-Enhanced Self-Training for Long-Form Personalized Text Generation
por: Salemi, Alireza, et al.
Publicado: (2025)
por: Salemi, Alireza, et al.
Publicado: (2025)
Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence
por: Ghasemabadi, Amirhosein, et al.
Publicado: (2025)
por: Ghasemabadi, Amirhosein, et al.
Publicado: (2025)
Ejemplares similares
-
Conformity in Large Language Models
por: Zhu, Xiaochen, et al.
Publicado: (2024) -
Atomic Calibration of LLMs in Long-Form Generations
por: Zhang, Caiqi, et al.
Publicado: (2024) -
All Roads Lead to Rome: Graph-Based Confidence Estimation for Large Language Model Reasoning
por: Zhang, Caiqi, et al.
Publicado: (2025) -
LoGU: Long-form Generation with Uncertainty Expressions
por: Yang, Ruihan, et al.
Publicado: (2024) -
Segment-Level Diffusion: A Framework for Controllable Long-Form Generation with Diffusion Language Models
por: Zhu, Xiaochen, et al.
Publicado: (2024)