CRAFT: Calibrated Reasoning with Answer-Faithful Traces via Reinforcement Learning for Multi-Hop Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yu, Zhang, Wenxiao, Guo, Diandian, Cao, Cong, Yuan, Fangfang, Sun, Qiang, Liu, Yanbing, Hong, Jin B., Ma, Zhiyuan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911517201399808
author Liu, Yu
Zhang, Wenxiao
Guo, Diandian
Cao, Cong
Yuan, Fangfang
Sun, Qiang
Liu, Yanbing
Hong, Jin B.
Ma, Zhiyuan
author_facet Liu, Yu
Zhang, Wenxiao
Guo, Diandian
Cao, Cong
Yuan, Fangfang
Sun, Qiang
Liu, Yanbing
Hong, Jin B.
Ma, Zhiyuan
contents Retrieval-augmented large language models, when optimized with outcome-level rewards, can achieve strong answer accuracy on multi-hop questions. However, under noisy retrieval, models frequently suffer from "right-answer-wrong-reason failures": they may exploit spurious shortcuts or produce reasoning traces weakly grounded in the supporting evidence. Furthermore, the lack of structured output control prevents reliable auditing of the underlying reasoning quality. To address this, we propose CRAFT (Calibrated Reasoning with Answer-Faithful Traces), a reinforcement learning framework for the response generation stage of retrieval-augmented multi-hop question answering. CRAFT trains models to produce structured reasoning traces with configurable levels of auditability (e.g., by selectively retaining planning, evidence citation, or reasoning steps). Training combines two complementary forms of supervision: deterministic rewards enforce verifiable constraints, including format compliance, answer correctness, and citation-set validity, while a judge-based reward audits semantic faithfulness by evaluating reasoning consistency and evidence grounding. Experiments show that CRAFT improves both answer accuracy and reasoning faithfulness across model scales. Notably, semantic judge-based rewards improve answer accuracy rather than compromise it, enabling CRAFT (7B) to achieve performance competitive with strong closed-source models.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01348
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CRAFT: Calibrated Reasoning with Answer-Faithful Traces via Reinforcement Learning for Multi-Hop Question Answering
Liu, Yu
Zhang, Wenxiao
Guo, Diandian
Cao, Cong
Yuan, Fangfang
Sun, Qiang
Liu, Yanbing
Hong, Jin B.
Ma, Zhiyuan
Computation and Language
Machine Learning
Retrieval-augmented large language models, when optimized with outcome-level rewards, can achieve strong answer accuracy on multi-hop questions. However, under noisy retrieval, models frequently suffer from "right-answer-wrong-reason failures": they may exploit spurious shortcuts or produce reasoning traces weakly grounded in the supporting evidence. Furthermore, the lack of structured output control prevents reliable auditing of the underlying reasoning quality. To address this, we propose CRAFT (Calibrated Reasoning with Answer-Faithful Traces), a reinforcement learning framework for the response generation stage of retrieval-augmented multi-hop question answering. CRAFT trains models to produce structured reasoning traces with configurable levels of auditability (e.g., by selectively retaining planning, evidence citation, or reasoning steps). Training combines two complementary forms of supervision: deterministic rewards enforce verifiable constraints, including format compliance, answer correctness, and citation-set validity, while a judge-based reward audits semantic faithfulness by evaluating reasoning consistency and evidence grounding. Experiments show that CRAFT improves both answer accuracy and reasoning faithfulness across model scales. Notably, semantic judge-based rewards improve answer accuracy rather than compromise it, enabling CRAFT (7B) to achieve performance competitive with strong closed-source models.
title CRAFT: Calibrated Reasoning with Answer-Faithful Traces via Reinforcement Learning for Multi-Hop Question Answering
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2602.01348