Accordion-Thinking: Self-Regulated Step Summaries for Efficient and Readable LLM Reasoning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yang, Zhicheng, Guo, Zhijiang, Huang, Yinya, Wang, Yongxin, Shi, Wenlei, Wang, Yiwei, Liang, Xiaodan, Tang, Jing
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915925694873600
author Yang, Zhicheng
Guo, Zhijiang
Huang, Yinya
Wang, Yongxin
Shi, Wenlei
Wang, Yiwei
Liang, Xiaodan
Tang, Jing
author_facet Yang, Zhicheng
Guo, Zhijiang
Huang, Yinya
Wang, Yongxin
Shi, Wenlei
Wang, Yiwei
Liang, Xiaodan
Tang, Jing
contents Scaling test-time compute via long Chain-of-Thought unlocks remarkable gains in reasoning capabilities, yet it faces practical limits due to the linear growth of KV cache and quadratic attention complexity. In this paper, we introduce Accordion-Thinking, an end-to-end framework where LLMs learn to self-regulate the granularity of the reasoning steps through dynamic summarization. This mechanism enables a Fold inference mode, where the model periodically summarizes its thought process and discards former thoughts to reduce dependency on historical tokens. We apply reinforcement learning to incentivize this capability further, uncovering a critical insight: the accuracy gap between the highly efficient Fold mode and the exhaustive Unfold mode progressively narrows and eventually vanishes over the course of training. This phenomenon demonstrates that the model learns to encode essential reasoning information into compact summaries, achieving effective compression of the reasoning context. Our Accordion-Thinking demonstrates that with learned self-compression, LLMs can tackle complex reasoning tasks with minimal dependency token overhead without compromising solution quality, and it achieves a three times throughput while maintaining accuracy on a 48GB GPU memory configuration, while the structured step summaries provide a human-readable account of the reasoning process.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03249
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Accordion-Thinking: Self-Regulated Step Summaries for Efficient and Readable LLM Reasoning
Yang, Zhicheng
Guo, Zhijiang
Huang, Yinya
Wang, Yongxin
Shi, Wenlei
Wang, Yiwei
Liang, Xiaodan
Tang, Jing
Artificial Intelligence
Machine Learning
Scaling test-time compute via long Chain-of-Thought unlocks remarkable gains in reasoning capabilities, yet it faces practical limits due to the linear growth of KV cache and quadratic attention complexity. In this paper, we introduce Accordion-Thinking, an end-to-end framework where LLMs learn to self-regulate the granularity of the reasoning steps through dynamic summarization. This mechanism enables a Fold inference mode, where the model periodically summarizes its thought process and discards former thoughts to reduce dependency on historical tokens. We apply reinforcement learning to incentivize this capability further, uncovering a critical insight: the accuracy gap between the highly efficient Fold mode and the exhaustive Unfold mode progressively narrows and eventually vanishes over the course of training. This phenomenon demonstrates that the model learns to encode essential reasoning information into compact summaries, achieving effective compression of the reasoning context. Our Accordion-Thinking demonstrates that with learned self-compression, LLMs can tackle complex reasoning tasks with minimal dependency token overhead without compromising solution quality, and it achieves a three times throughput while maintaining accuracy on a 48GB GPU memory configuration, while the structured step summaries provide a human-readable account of the reasoning process.
title Accordion-Thinking: Self-Regulated Step Summaries for Efficient and Readable LLM Reasoning
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2602.03249