AutoL2S: Auto Long-Short Reasoning for Efficient Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Feng, Chuang, Yu-Neng, Wang, Guanchu, Le, Hoang Anh Duy, Zhong, Shaochen, Liu, Hongyi, Yuan, Jiayi, Sui, Yang, Braverman, Vladimir, Chaudhary, Vipin, Hu, Xia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909983976718336
author Luo, Feng
Chuang, Yu-Neng
Wang, Guanchu
Le, Hoang Anh Duy
Zhong, Shaochen
Liu, Hongyi
Yuan, Jiayi
Sui, Yang
Braverman, Vladimir
Chaudhary, Vipin
Hu, Xia
author_facet Luo, Feng
Chuang, Yu-Neng
Wang, Guanchu
Le, Hoang Anh Duy
Zhong, Shaochen
Liu, Hongyi
Yuan, Jiayi
Sui, Yang
Braverman, Vladimir
Chaudhary, Vipin
Hu, Xia
contents Reasoning-capable large language models (LLMs) achieve strong performance on complex tasks but often exhibit overthinking after distillation, generating unnecessarily long chain-of-thought (CoT) reasoning even for simple inputs and incurring high inference cost. However, naively shortening reasoning length can degrade reasoning accuracy, as concise reasoning may be insufficient for certain inputs and lacks explicit supervision. We propose Auto Long-Short Reasoning (AutoL2S), a distillation framework that empowers non-reasoning LLMs to think thoroughly but only when necessary. AutoL2S first learns a lightweight switching token with verified long-short CoTs to enable instance-wise long-short reasoning selection. Then it leverages long-short reasoning rollouts induced by a switching token in a GRPO-style loss to improve reasoning efficiency while maintaining accuracy. Experiments demonstrate that AutoL2S effectively reduces reasoning length up to 71% with minimal accuracy loss, yielding markedly better trade-off in token length and inference time while preserving accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22662
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AutoL2S: Auto Long-Short Reasoning for Efficient Large Language Models
Luo, Feng
Chuang, Yu-Neng
Wang, Guanchu
Le, Hoang Anh Duy
Zhong, Shaochen
Liu, Hongyi
Yuan, Jiayi
Sui, Yang
Braverman, Vladimir
Chaudhary, Vipin
Hu, Xia
Computation and Language
Machine Learning
Reasoning-capable large language models (LLMs) achieve strong performance on complex tasks but often exhibit overthinking after distillation, generating unnecessarily long chain-of-thought (CoT) reasoning even for simple inputs and incurring high inference cost. However, naively shortening reasoning length can degrade reasoning accuracy, as concise reasoning may be insufficient for certain inputs and lacks explicit supervision. We propose Auto Long-Short Reasoning (AutoL2S), a distillation framework that empowers non-reasoning LLMs to think thoroughly but only when necessary. AutoL2S first learns a lightweight switching token with verified long-short CoTs to enable instance-wise long-short reasoning selection. Then it leverages long-short reasoning rollouts induced by a switching token in a GRPO-style loss to improve reasoning efficiency while maintaining accuracy. Experiments demonstrate that AutoL2S effectively reduces reasoning length up to 71% with minimal accuracy loss, yielding markedly better trade-off in token length and inference time while preserving accuracy.
title AutoL2S: Auto Long-Short Reasoning for Efficient Large Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.22662