Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Jiwon, Jo, Dongwon, Kim, Yulhwa, Kim, Jae-Joon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
by: Jo, Dongwon, et al.
Published: (2025)
by: Jo, Dongwon, et al.
Published: (2025)
RelayGen: Intra-Generation Model Switching for Efficient Reasoning
by: Song, Jiwon, et al.
Published: (2026)
by: Song, Jiwon, et al.
Published: (2026)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
by: Jo, Dongwon, et al.
Published: (2026)
by: Jo, Dongwon, et al.
Published: (2026)
CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection
by: Song, Jiwon, et al.
Published: (2026)
by: Song, Jiwon, et al.
Published: (2026)
SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks
by: Song, Jiwon, et al.
Published: (2024)
by: Song, Jiwon, et al.
Published: (2024)
LiteStage: Latency-aware Layer Skipping for Multi-stage Reasoning
by: Kang, Beomseok, et al.
Published: (2025)
by: Kang, Beomseok, et al.
Published: (2025)
Mixture of Scales: Memory-Efficient Token-Adaptive Binarization for Large Language Models
by: Jo, Dongwon, et al.
Published: (2024)
by: Jo, Dongwon, et al.
Published: (2024)
QWHA: Quantization-Aware Walsh-Hadamard Adaptation for Parameter-Efficient Fine-Tuning on Large Language Models
by: Jeon, Hyesung, et al.
Published: (2025)
by: Jeon, Hyesung, et al.
Published: (2025)
L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language Models
by: Jeon, Hyesung, et al.
Published: (2024)
by: Jeon, Hyesung, et al.
Published: (2024)
Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference
by: Kang, Beomseok, et al.
Published: (2026)
by: Kang, Beomseok, et al.
Published: (2026)
Breadcrumbs Reasoning: Memory-Efficient Reasoning with Compression Beacons
by: Monea, Giovanni, et al.
Published: (2025)
by: Monea, Giovanni, et al.
Published: (2025)
Think Less, Know More: State-Aware Reasoning Compression with Knowledge Guidance for Efficient Reasoning
by: Sui, Yi, et al.
Published: (2026)
by: Sui, Yi, et al.
Published: (2026)
Code Execution as Grounded Supervision for LLM Reasoning
by: Jung, Dongwon, et al.
Published: (2025)
by: Jung, Dongwon, et al.
Published: (2025)
PACE: Prefix-Protected and Difficulty-Aware Compression for Efficient Reasoning
by: Feng, Ruixiang, et al.
Published: (2026)
by: Feng, Ruixiang, et al.
Published: (2026)
Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains
by: Tan, Wenhui, et al.
Published: (2025)
by: Tan, Wenhui, et al.
Published: (2025)
TriAttention: Efficient Long Reasoning with Trigonometric KV Compression
by: Mao, Weian, et al.
Published: (2026)
by: Mao, Weian, et al.
Published: (2026)
Selective Latent Thinking: Adaptive Compression of LLM Reasoning Chains
by: Xie, Hui, et al.
Published: (2026)
by: Xie, Hui, et al.
Published: (2026)
Entropy-Guided Reasoning Compression
by: Zhu, Hourun, et al.
Published: (2025)
by: Zhu, Hourun, et al.
Published: (2025)
GRC: Unifying Reasoning-Driven Generation, Retrieval and Compression
by: Miao, Zhongtao, et al.
Published: (2026)
by: Miao, Zhongtao, et al.
Published: (2026)
Compressed Chain of Thought: Efficient Reasoning Through Dense Representations
by: Cheng, Jeffrey, et al.
Published: (2024)
by: Cheng, Jeffrey, et al.
Published: (2024)
Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens
by: Zhao, Zhenyu, et al.
Published: (2026)
by: Zhao, Zhenyu, et al.
Published: (2026)
OPSD Compresses What RLVR Teaches: A Post-RL Compaction Stage for Reasoning Models
by: Kim, Jaehoon, et al.
Published: (2026)
by: Kim, Jaehoon, et al.
Published: (2026)
TL;DR: Too Long, Do Re-weighting for Efficient LLM Reasoning Compression
by: Li, Zhong-Zhi, et al.
Published: (2025)
by: Li, Zhong-Zhi, et al.
Published: (2025)
The Path of Least Resistance: Guiding LLM Reasoning Trajectories with Prefix Consensus
by: Jindal, Ishan, et al.
Published: (2026)
by: Jindal, Ishan, et al.
Published: (2026)
LongFlow: Efficient KV Cache Compression for Reasoning Models
by: Su, Yi, et al.
Published: (2026)
by: Su, Yi, et al.
Published: (2026)
R-Capsule: Compressing High-Level Plans for Efficient Large Language Model Reasoning
by: Shan, Hongyu, et al.
Published: (2025)
by: Shan, Hongyu, et al.
Published: (2025)
Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring
by: Zhang, Dongxu, et al.
Published: (2026)
by: Zhang, Dongxu, et al.
Published: (2026)
Reasoning Aware Self-Consistency: Leveraging Reasoning Paths for Efficient LLM Sampling
by: Wan, Guangya, et al.
Published: (2024)
by: Wan, Guangya, et al.
Published: (2024)
DSPC: Dual-Stage Progressive Compression Framework for Efficient Long-Context Reasoning
by: Gao, Yaxin, et al.
Published: (2025)
by: Gao, Yaxin, et al.
Published: (2025)
A*-Thought: Efficient Reasoning via Bidirectional Compression for Low-Resource Settings
by: Xu, Xiaoang, et al.
Published: (2025)
by: Xu, Xiaoang, et al.
Published: (2025)
VTC-R1: Vision-Text Compression for Efficient Long-Context Reasoning
by: Wang, Yibo, et al.
Published: (2026)
by: Wang, Yibo, et al.
Published: (2026)
ConMax: Confidence-Maximizing Compression for Efficient Chain-of-Thought Reasoning
by: Hu, Minda, et al.
Published: (2026)
by: Hu, Minda, et al.
Published: (2026)
Familiarity-Aware Evidence Compression for Retrieval-Augmented Generation
by: Jung, Dongwon, et al.
Published: (2024)
by: Jung, Dongwon, et al.
Published: (2024)
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
by: Kim, Jeonghye, et al.
Published: (2026)
by: Kim, Jeonghye, et al.
Published: (2026)
Guiding Reasoning in Small Language Models with LLM Assistance
by: Kim, Yujin, et al.
Published: (2025)
by: Kim, Yujin, et al.
Published: (2025)
Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
by: Liu, Minghui, et al.
Published: (2025)
by: Liu, Minghui, et al.
Published: (2025)
ThinkBrake: Efficient Reasoning via Log-Probability Margin Guided Decoding
by: Song, Sangjun, et al.
Published: (2025)
by: Song, Sangjun, et al.
Published: (2025)
TensorLLM: Tensorising Multi-Head Attention for Enhanced Reasoning and Compression in LLMs
by: Gu, Yuxuan, et al.
Published: (2025)
by: Gu, Yuxuan, et al.
Published: (2025)
Optimizing Length Compression in Large Reasoning Models
by: Cheng, Zhengxiang, et al.
Published: (2025)
by: Cheng, Zhengxiang, et al.
Published: (2025)
CORRECT: Context- and Reference-Augmented Reasoning and Prompting for Fact-Checking
by: Zhang, Delvin Ce, et al.
Published: (2025)
by: Zhang, Delvin Ce, et al.
Published: (2025)
Similar Items
-
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
by: Jo, Dongwon, et al.
Published: (2025) -
RelayGen: Intra-Generation Model Switching for Efficient Reasoning
by: Song, Jiwon, et al.
Published: (2026) -
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
by: Jo, Dongwon, et al.
Published: (2026) -
CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection
by: Song, Jiwon, et al.
Published: (2026) -
SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks
by: Song, Jiwon, et al.
Published: (2024)