FRIT: Using Causal Importance to Improve Chain-of-Thought Faithfulness

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Swaroop, Anand, Nallani, Akshat, Uboweja, Saksham, Uzdenova, Adiliia, Nguyen, Michael, Zhu, Kevin, Dev, Sunishchal, Panda, Ashwinee, Sharma, Vasu, Chaudhary, Maheep
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916952823300096
author Swaroop, Anand
Nallani, Akshat
Uboweja, Saksham
Uzdenova, Adiliia
Nguyen, Michael
Zhu, Kevin
Dev, Sunishchal
Panda, Ashwinee
Sharma, Vasu
Chaudhary, Maheep
author_facet Swaroop, Anand
Nallani, Akshat
Uboweja, Saksham
Uzdenova, Adiliia
Nguyen, Michael
Zhu, Kevin
Dev, Sunishchal
Panda, Ashwinee
Sharma, Vasu
Chaudhary, Maheep
contents Chain-of-thought (CoT) reasoning has emerged as a powerful tool for improving large language model performance on complex tasks, but recent work shows that reasoning steps often fail to causally influence the final answer, creating brittle and untrustworthy outputs. Prior approaches focus primarily on measuring faithfulness, while methods for systematically improving it remain limited. We introduce Faithful Reasoning via Intervention Training (FRIT), a scalable alignment method that trains models to produce causally consistent reasoning by learning from systematically corrupted examples. FRIT generates synthetic training data by intervening on individual reasoning steps in model-generated CoTs, creating faithful/unfaithful pairs that highlight when reasoning breaks down. We then apply Direct Preference Optimization to teach models to prefer causally consistent reasoning paths. Evaluating on Qwen3-8B and Mistral-7B-v0.1 across factual and symbolic reasoning tasks, FRIT increases faithful reasoning by $3.4$ percentage points for Mistral on GSM8K while improving accuracy by $7.6$ percentage points. Our approach provides the first scalable, supervision-free method for training language models to produce more reliable and interpretable reasoning, addressing a critical gap between reasoning performance and trustworthiness. We release our code at \href{https://github.com/Anut-py/frit}.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13334
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FRIT: Using Causal Importance to Improve Chain-of-Thought Faithfulness
Swaroop, Anand
Nallani, Akshat
Uboweja, Saksham
Uzdenova, Adiliia
Nguyen, Michael
Zhu, Kevin
Dev, Sunishchal
Panda, Ashwinee
Sharma, Vasu
Chaudhary, Maheep
Artificial Intelligence
Machine Learning
Chain-of-thought (CoT) reasoning has emerged as a powerful tool for improving large language model performance on complex tasks, but recent work shows that reasoning steps often fail to causally influence the final answer, creating brittle and untrustworthy outputs. Prior approaches focus primarily on measuring faithfulness, while methods for systematically improving it remain limited. We introduce Faithful Reasoning via Intervention Training (FRIT), a scalable alignment method that trains models to produce causally consistent reasoning by learning from systematically corrupted examples. FRIT generates synthetic training data by intervening on individual reasoning steps in model-generated CoTs, creating faithful/unfaithful pairs that highlight when reasoning breaks down. We then apply Direct Preference Optimization to teach models to prefer causally consistent reasoning paths. Evaluating on Qwen3-8B and Mistral-7B-v0.1 across factual and symbolic reasoning tasks, FRIT increases faithful reasoning by $3.4$ percentage points for Mistral on GSM8K while improving accuracy by $7.6$ percentage points. Our approach provides the first scalable, supervision-free method for training language models to produce more reliable and interpretable reasoning, addressing a critical gap between reasoning performance and trustworthiness. We release our code at \href{https://github.com/Anut-py/frit}.
title FRIT: Using Causal Importance to Improve Chain-of-Thought Faithfulness
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.13334