TL;DR: Too Long, Do Re-weighting for Efficient LLM Reasoning Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhong-Zhi, Liang, Xiao, Tang, Zihao, Ji, Lei, Wang, Peijie, Xu, Haotian, W, Xing, Huang, Haizhen, Deng, Weiwei, Gong, Yeyun, Guo, Zhijiang, Liu, Xiao, Yin, Fei, Liu, Cheng-Lin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908407647174656
author Li, Zhong-Zhi
Liang, Xiao
Tang, Zihao
Ji, Lei
Wang, Peijie
Xu, Haotian
W, Xing
Huang, Haizhen
Deng, Weiwei
Gong, Yeyun
Guo, Zhijiang
Liu, Xiao
Yin, Fei
Liu, Cheng-Lin
author_facet Li, Zhong-Zhi
Liang, Xiao
Tang, Zihao
Ji, Lei
Wang, Peijie
Xu, Haotian
W, Xing
Huang, Haizhen
Deng, Weiwei
Gong, Yeyun
Guo, Zhijiang
Liu, Xiao
Yin, Fei
Liu, Cheng-Lin
contents Large Language Models (LLMs) have recently achieved remarkable progress by leveraging Reinforcement Learning and extended Chain-of-Thought (CoT) techniques. However, the challenge of performing efficient language reasoning--especially during inference with extremely long outputs--has drawn increasing attention from the research community. In this work, we propose a dynamic ratio-based training pipeline that does not rely on sophisticated data annotations or interpolation between multiple models. We continuously balance the weights between the model's System-1 and System-2 data to eliminate redundant reasoning processes while preserving the model's reasoning capability. We validate our approach across models on DeepSeek-R1-Distill-7B and DeepSeek-R1-Distill-14B and on a diverse set of benchmarks with varying difficulty levels. Our method significantly reduces the number of output tokens by nearly 40% while maintaining the accuracy of the reasoning. Our code and data will be available soon.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02678
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TL;DR: Too Long, Do Re-weighting for Efficient LLM Reasoning Compression
Li, Zhong-Zhi
Liang, Xiao
Tang, Zihao
Ji, Lei
Wang, Peijie
Xu, Haotian
W, Xing
Huang, Haizhen
Deng, Weiwei
Gong, Yeyun
Guo, Zhijiang
Liu, Xiao
Yin, Fei
Liu, Cheng-Lin
Computation and Language
Computational Engineering, Finance, and Science
Numerical Analysis
Large Language Models (LLMs) have recently achieved remarkable progress by leveraging Reinforcement Learning and extended Chain-of-Thought (CoT) techniques. However, the challenge of performing efficient language reasoning--especially during inference with extremely long outputs--has drawn increasing attention from the research community. In this work, we propose a dynamic ratio-based training pipeline that does not rely on sophisticated data annotations or interpolation between multiple models. We continuously balance the weights between the model's System-1 and System-2 data to eliminate redundant reasoning processes while preserving the model's reasoning capability. We validate our approach across models on DeepSeek-R1-Distill-7B and DeepSeek-R1-Distill-14B and on a diverse set of benchmarks with varying difficulty levels. Our method significantly reduces the number of output tokens by nearly 40% while maintaining the accuracy of the reasoning. Our code and data will be available soon.
title TL;DR: Too Long, Do Re-weighting for Efficient LLM Reasoning Compression
topic Computation and Language
Computational Engineering, Finance, and Science
Numerical Analysis
url https://arxiv.org/abs/2506.02678