Scalable LLM Reasoning Acceleration with Low-rank Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dong, Harry, Acun, Bilge, Chen, Beidi, Chi, Yuejie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908834902048768
author Dong, Harry
Acun, Bilge
Chen, Beidi
Chi, Yuejie
author_facet Dong, Harry
Acun, Bilge
Chen, Beidi
Chi, Yuejie
contents Due to long generations, large language model (LLM) math reasoning demands significant computational resources and time. While many existing efficient inference methods have been developed with excellent performance preservation on language tasks, they often severely degrade math performance. In this paper, we propose Caprese, a resource-efficient distillation method to recover lost capabilities from deploying efficient inference methods, focused primarily in feedforward blocks. With original weights unperturbed, roughly 1% of additional parameters, and only 20K synthetic training samples, we are able to recover much if not all of the reasoning capabilities lost from efficient inference for thinking LLMs and without harm to language tasks for instruct LLMs. Moreover, Caprese slashes the number of active parameters (~2B cut for Gemma 2 9B and Llama 3.1 8B) and integrates cleanly into existing model layers to reduce latency (>16% time-to-next-token reduction) while encouraging response brevity (up to 8.5% fewer tokens).
format Preprint
id arxiv_https___arxiv_org_abs_2505_07861
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scalable LLM Reasoning Acceleration with Low-rank Distillation
Dong, Harry
Acun, Bilge
Chen, Beidi
Chi, Yuejie
Computation and Language
Artificial Intelligence
Machine Learning
Due to long generations, large language model (LLM) math reasoning demands significant computational resources and time. While many existing efficient inference methods have been developed with excellent performance preservation on language tasks, they often severely degrade math performance. In this paper, we propose Caprese, a resource-efficient distillation method to recover lost capabilities from deploying efficient inference methods, focused primarily in feedforward blocks. With original weights unperturbed, roughly 1% of additional parameters, and only 20K synthetic training samples, we are able to recover much if not all of the reasoning capabilities lost from efficient inference for thinking LLMs and without harm to language tasks for instruct LLMs. Moreover, Caprese slashes the number of active parameters (~2B cut for Gemma 2 9B and Llama 3.1 8B) and integrates cleanly into existing model layers to reduce latency (>16% time-to-next-token reduction) while encouraging response brevity (up to 8.5% fewer tokens).
title Scalable LLM Reasoning Acceleration with Low-rank Distillation
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.07861