Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Guo, Zhenyuan, Chen, Tong, Meng, Wenlong, Gong, Chen, Yu, Xin, Wei, Chengkun, Chen, Wenzhi
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912849598611456
author Guo, Zhenyuan
Chen, Tong
Meng, Wenlong
Gong, Chen
Yu, Xin
Wei, Chengkun
Chen, Wenzhi
author_facet Guo, Zhenyuan
Chen, Tong
Meng, Wenlong
Gong, Chen
Yu, Xin
Wei, Chengkun
Chen, Wenzhi
contents Large Reasoning Models (LRMs) excel at solving complex problems by explicitly generating a reasoning trace before deriving the final answer. However, these extended generations incur substantial memory footprint and computational overhead, bottlenecking LRMs' efficiency. This work uses attention maps to analyze the influence of reasoning traces and uncover an interesting phenomenon: only some decision-critical tokens in a reasoning trace steer the model toward the final answer, while the remaining tokens contribute negligibly. Building on this observation, we propose Dynamic Thinking-Token Selection (DynTS). This method identifies decision-critical tokens and retains only their associated Key-Value (KV) cache states during inference, evicting the remaining redundant entries to optimize efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2601_18383
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Models
Guo, Zhenyuan
Chen, Tong
Meng, Wenlong
Gong, Chen
Yu, Xin
Wei, Chengkun
Chen, Wenzhi
Artificial Intelligence
Computation and Language
Machine Learning
Large Reasoning Models (LRMs) excel at solving complex problems by explicitly generating a reasoning trace before deriving the final answer. However, these extended generations incur substantial memory footprint and computational overhead, bottlenecking LRMs' efficiency. This work uses attention maps to analyze the influence of reasoning traces and uncover an interesting phenomenon: only some decision-critical tokens in a reasoning trace steer the model toward the final answer, while the remaining tokens contribute negligibly. Building on this observation, we propose Dynamic Thinking-Token Selection (DynTS). This method identifies decision-critical tokens and retains only their associated Key-Value (KV) cache states during inference, evicting the remaining redundant entries to optimize efficiency.
title Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Models
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2601.18383