Efficient Inference for Large Reasoning Models: A Survey

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yue, Wu, Jiaying, He, Yufei, Gong, Ruihan, Xia, Jun, Li, Liang, Gao, Hongcheng, Chen, Hongyu, Bi, Baolong, Zhang, Jiaheng, Huang, Zhiqi, Hooi, Bryan, Li, Stan Z., Li, Keqin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915444633370624
author Liu, Yue
Wu, Jiaying
He, Yufei
Gong, Ruihan
Xia, Jun
Li, Liang
Gao, Hongcheng
Chen, Hongyu
Bi, Baolong
Zhang, Jiaheng
Huang, Zhiqi
Hooi, Bryan
Li, Stan Z.
Li, Keqin
author_facet Liu, Yue
Wu, Jiaying
He, Yufei
Gong, Ruihan
Xia, Jun
Li, Liang
Gao, Hongcheng
Chen, Hongyu
Bi, Baolong
Zhang, Jiaheng
Huang, Zhiqi
Hooi, Bryan
Li, Stan Z.
Li, Keqin
contents Large Reasoning Models (LRMs) significantly improve the reasoning ability of Large Language Models (LLMs) by learning to reason, exhibiting promising performance in solving complex tasks. However, their deliberative reasoning process leads to inefficiencies in token usage, memory consumption, and inference time. Thus, this survey provides a review of efficient inference methods designed specifically for LRMs, focusing on mitigating token inefficiency while preserving the reasoning quality. The overview structure of this paper is shown in Figure~\ref{fig:paper_structure}. First, we introduce a taxonomy to group the recent methods into two main categories: (a) explicit compact Chain-of-Thought (CoT), which reduces tokens while keeping the explicit reasoning structure, and (b) implicit latent CoT, which encodes reasoning steps within hidden representations instead of explicit tokens. Meanwhile, we discuss their strengths and weaknesses. Then, we conduct empirical analyses on existing methods from reasoning scenarios, object functions, and performance \& efficiency aspects. Besides, we present open challenges in this field, including human-centric controllable reasoning, trade-off between interpretability and efficiency of reasoning, ensuring the safety of efficient reasoning, and broader applications of efficient reasoning. In addition, we highlight key insights for enhancing LRMs' inference efficiency via techniques such as model merging, new architectures, and agent routers. We hope this work serves as a valuable guide, helping researchers overcome challenges in this vibrant field. A collection of efficient reasoning methods for LRMs (papers and codes) is provided at this link: https://github.com/yueliu1999/Awesome-Efficient-Inference-for-LRMs.
format Preprint
id arxiv_https___arxiv_org_abs_2503_23077
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Inference for Large Reasoning Models: A Survey
Liu, Yue
Wu, Jiaying
He, Yufei
Gong, Ruihan
Xia, Jun
Li, Liang
Gao, Hongcheng
Chen, Hongyu
Bi, Baolong
Zhang, Jiaheng
Huang, Zhiqi
Hooi, Bryan
Li, Stan Z.
Li, Keqin
Computation and Language
Large Reasoning Models (LRMs) significantly improve the reasoning ability of Large Language Models (LLMs) by learning to reason, exhibiting promising performance in solving complex tasks. However, their deliberative reasoning process leads to inefficiencies in token usage, memory consumption, and inference time. Thus, this survey provides a review of efficient inference methods designed specifically for LRMs, focusing on mitigating token inefficiency while preserving the reasoning quality. The overview structure of this paper is shown in Figure~\ref{fig:paper_structure}. First, we introduce a taxonomy to group the recent methods into two main categories: (a) explicit compact Chain-of-Thought (CoT), which reduces tokens while keeping the explicit reasoning structure, and (b) implicit latent CoT, which encodes reasoning steps within hidden representations instead of explicit tokens. Meanwhile, we discuss their strengths and weaknesses. Then, we conduct empirical analyses on existing methods from reasoning scenarios, object functions, and performance \& efficiency aspects. Besides, we present open challenges in this field, including human-centric controllable reasoning, trade-off between interpretability and efficiency of reasoning, ensuring the safety of efficient reasoning, and broader applications of efficient reasoning. In addition, we highlight key insights for enhancing LRMs' inference efficiency via techniques such as model merging, new architectures, and agent routers. We hope this work serves as a valuable guide, helping researchers overcome challenges in this vibrant field. A collection of efficient reasoning methods for LRMs (papers and codes) is provided at this link: https://github.com/yueliu1999/Awesome-Efficient-Inference-for-LRMs.
title Efficient Inference for Large Reasoning Models: A Survey
topic Computation and Language
url https://arxiv.org/abs/2503.23077