Think Clearly: Improving Reasoning via Redundant Token Pruning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Choi, Daewon, Lee, Jimin, Tack, Jihoon, Song, Woomin, Dingliwal, Saket, Jayanthi, Sai Muralidhar, Ganesh, Bhavana, Shin, Jinwoo, Galstyan, Aram, Bodapati, Sravan Babu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908447112429568
author Choi, Daewon
Lee, Jimin
Tack, Jihoon
Song, Woomin
Dingliwal, Saket
Jayanthi, Sai Muralidhar
Ganesh, Bhavana
Shin, Jinwoo
Galstyan, Aram
Bodapati, Sravan Babu
author_facet Choi, Daewon
Lee, Jimin
Tack, Jihoon
Song, Woomin
Dingliwal, Saket
Jayanthi, Sai Muralidhar
Ganesh, Bhavana
Shin, Jinwoo
Galstyan, Aram
Bodapati, Sravan Babu
contents Recent large language models have shown promising capabilities in long-form reasoning, following structured chains of thought before arriving at a final answer. However, we observe that these reasoning paths tend to include substantial redundancy; analyzing attention patterns reveals that attention scores are widely scattered, particularly incorrect answers exhibit greater attention sparsity. In this paper, we demonstrate that deliberately removing this redundancy in the reasoning process significantly improves performance through clear thinking, i.e., removing distraction. Specifically, we systematically identify reasoning redundancy by measuring token-level attention scores to a special end-of-thinking token, which is appended to an explicit instruction inserted to conclude each intermediate reasoning step. Furthermore, we propose structure-aware pruning that prioritizes removing tokens in low-contributing reasoning chunks over individual tokens. After evicting redundant tokens, we remove the injected end-of-thinking instruction, then resume the reasoning generation. We demonstrate that our method significantly improves overall accuracy across reasoning-intensive benchmarks without any training involved. In particular, our method shows strong performance on challenging mathematical competition benchmarks such as AIME and AMC, where reasoning redundancy is more prevalent.
format Preprint
id arxiv_https___arxiv_org_abs_2507_08806
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Think Clearly: Improving Reasoning via Redundant Token Pruning
Choi, Daewon
Lee, Jimin
Tack, Jihoon
Song, Woomin
Dingliwal, Saket
Jayanthi, Sai Muralidhar
Ganesh, Bhavana
Shin, Jinwoo
Galstyan, Aram
Bodapati, Sravan Babu
Artificial Intelligence
Computation and Language
Machine Learning
Recent large language models have shown promising capabilities in long-form reasoning, following structured chains of thought before arriving at a final answer. However, we observe that these reasoning paths tend to include substantial redundancy; analyzing attention patterns reveals that attention scores are widely scattered, particularly incorrect answers exhibit greater attention sparsity. In this paper, we demonstrate that deliberately removing this redundancy in the reasoning process significantly improves performance through clear thinking, i.e., removing distraction. Specifically, we systematically identify reasoning redundancy by measuring token-level attention scores to a special end-of-thinking token, which is appended to an explicit instruction inserted to conclude each intermediate reasoning step. Furthermore, we propose structure-aware pruning that prioritizes removing tokens in low-contributing reasoning chunks over individual tokens. After evicting redundant tokens, we remove the injected end-of-thinking instruction, then resume the reasoning generation. We demonstrate that our method significantly improves overall accuracy across reasoning-intensive benchmarks without any training involved. In particular, our method shows strong performance on challenging mathematical competition benchmarks such as AIME and AMC, where reasoning redundancy is more prevalent.
title Think Clearly: Improving Reasoning via Redundant Token Pruning
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2507.08806