Saved in:
Bibliographic Details
Main Authors: Wang, Xinming, Xu, Jian, Yu, Bin, Lian, Sheng, Yi, Hongzhu, Chen, Yi, Zhu, Yingjian, Wang, Boran, Yang, Hongming, Hu, Han, Zhang, Xu-Yao, Liu, Cheng-Lin
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.24794
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912801293860864
author Wang, Xinming
Xu, Jian
Yu, Bin
Lian, Sheng
Yi, Hongzhu
Chen, Yi
Zhu, Yingjian
Wang, Boran
Yang, Hongming
Hu, Han
Zhang, Xu-Yao
Liu, Cheng-Lin
author_facet Wang, Xinming
Xu, Jian
Yu, Bin
Lian, Sheng
Yi, Hongzhu
Chen, Yi
Zhu, Yingjian
Wang, Boran
Yang, Hongming
Hu, Han
Zhang, Xu-Yao
Liu, Cheng-Lin
contents Large reasoning models (LRMs) show strong capabilities in complex reasoning, yet their marginal gains on evidence-dependent factual questions are limited. We find this limitation is partially attributable to a reasoning-answer hit gap, where the model identifies the correct facts during reasoning but fails to incorporate them into the final response, thereby reducing factual fidelity. To address this issue, we propose MR-ALIGN, a Meta-Reasoning informed alignment framework that enhances factuality without relying on external verifiers. MR-ALIGN quantifies state transition probabilities along the model's thinking process and constructs a transition-aware implicit reward that reinforces beneficial reasoning patterns while suppressing defective ones at the atomic thinking segments. This re-weighting reshapes token-level signals into probability-aware segment scores, encouraging coherent reasoning trajectories that are more conducive to factual correctness. Empirical evaluations across four factual QA datasets and one long-form factuality benchmark show that MR-ALIGN consistently improves accuracy and truthfulness while reducing misleading reasoning. These results highlight that aligning the reasoning process itself, rather than merely the outputs, is pivotal for advancing factuality in LRMs.
format Preprint
id arxiv_https___arxiv_org_abs_2510_24794
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MR-Align: Meta-Reasoning Informed Factuality Alignment for Large Reasoning Models
Wang, Xinming
Xu, Jian
Yu, Bin
Lian, Sheng
Yi, Hongzhu
Chen, Yi
Zhu, Yingjian
Wang, Boran
Yang, Hongming
Hu, Han
Zhang, Xu-Yao
Liu, Cheng-Lin
Computation and Language
Large reasoning models (LRMs) show strong capabilities in complex reasoning, yet their marginal gains on evidence-dependent factual questions are limited. We find this limitation is partially attributable to a reasoning-answer hit gap, where the model identifies the correct facts during reasoning but fails to incorporate them into the final response, thereby reducing factual fidelity. To address this issue, we propose MR-ALIGN, a Meta-Reasoning informed alignment framework that enhances factuality without relying on external verifiers. MR-ALIGN quantifies state transition probabilities along the model's thinking process and constructs a transition-aware implicit reward that reinforces beneficial reasoning patterns while suppressing defective ones at the atomic thinking segments. This re-weighting reshapes token-level signals into probability-aware segment scores, encouraging coherent reasoning trajectories that are more conducive to factual correctness. Empirical evaluations across four factual QA datasets and one long-form factuality benchmark show that MR-ALIGN consistently improves accuracy and truthfulness while reducing misleading reasoning. These results highlight that aligning the reasoning process itself, rather than merely the outputs, is pivotal for advancing factuality in LRMs.
title MR-Align: Meta-Reasoning Informed Factuality Alignment for Large Reasoning Models
topic Computation and Language
url https://arxiv.org/abs/2510.24794