AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Höth, Max Henning, Kersting, Kristian, Deiseroth, Björn, Parcalabescu, Letitia
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917416148140032
author Höth, Max Henning
Kersting, Kristian
Deiseroth, Björn
Parcalabescu, Letitia
author_facet Höth, Max Henning
Kersting, Kristian
Deiseroth, Björn
Parcalabescu, Letitia
contents Large language models (LLMs) increasingly rely on chain-of-thought (CoT) reasoning to solve complex tasks. Yet ensuring that the reasoning trace both contributes to and faithfully reflects the processes underlying the model's final answer, rather than merely accompanying it, remains challenging. We introduce AtManRL, a method that leverages differentiable attention manipulation to learn more faithful reasoning through reinforcement learning. By training an additive attention mask that identifies tokens in the CoT crucial for producing correct answers, we derive a saliency reward signal that encourages the model to generate reasoning traces that genuinely influence its final predictions. We integrate this saliency reward with outcome-based rewards within the GRPO framework to jointly optimize for correctness and interpretability. Experiments on GSM8K and MMLU with Llama-3.2-3B-Instruct demonstrate that our approach can identify influential reasoning tokens and enable training more transparent reasoning models.
format Preprint
id arxiv_https___arxiv_org_abs_2604_16158
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency
Höth, Max Henning
Kersting, Kristian
Deiseroth, Björn
Parcalabescu, Letitia
Computation and Language
Artificial Intelligence
Machine Learning
I.2.7
Large language models (LLMs) increasingly rely on chain-of-thought (CoT) reasoning to solve complex tasks. Yet ensuring that the reasoning trace both contributes to and faithfully reflects the processes underlying the model's final answer, rather than merely accompanying it, remains challenging. We introduce AtManRL, a method that leverages differentiable attention manipulation to learn more faithful reasoning through reinforcement learning. By training an additive attention mask that identifies tokens in the CoT crucial for producing correct answers, we derive a saliency reward signal that encourages the model to generate reasoning traces that genuinely influence its final predictions. We integrate this saliency reward with outcome-based rewards within the GRPO framework to jointly optimize for correctness and interpretability. Experiments on GSM8K and MMLU with Llama-3.2-3B-Instruct demonstrate that our approach can identify influential reasoning tokens and enable training more transparent reasoning models.
title AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency
topic Computation and Language
Artificial Intelligence
Machine Learning
I.2.7
url https://arxiv.org/abs/2604.16158