Token-Budget-Aware LLM Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Tingxu, Wang, Zhenting, Fang, Chunrong, Zhao, Shiyu, Ma, Shiqing, Chen, Zhenyu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912407448715264
author Han, Tingxu
Wang, Zhenting
Fang, Chunrong
Zhao, Shiyu
Ma, Shiqing
Chen, Zhenyu
author_facet Han, Tingxu
Wang, Zhenting
Fang, Chunrong
Zhao, Shiyu
Ma, Shiqing
Chen, Zhenyu
contents Reasoning is critical for large language models (LLMs) to excel in a wide range of tasks. While methods like Chain-of-Thought (CoT) reasoning and enhance LLM performance by decomposing problems into intermediate steps, they also incur significant overhead in token usage, leading to increased costs. We find that the reasoning process of current LLMs is unnecessarily lengthy and it can be compressed by including a reasonable token budget in the prompt, but the choice of token budget plays a crucial role in the actual compression effectiveness. We then propose a token-budget-aware LLM reasoning framework that dynamically adjusts the number of reasoning tokens based on the reasoning complexity of each problem. Experiments show that our method effectively reduces token costs in CoT reasoning with only a slight performance reduction, offering a practical solution to balance efficiency and accuracy in LLM reasoning. Code: https://github.com/GeniusHTX/TALE
format Preprint
id arxiv_https___arxiv_org_abs_2412_18547
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Token-Budget-Aware LLM Reasoning
Han, Tingxu
Wang, Zhenting
Fang, Chunrong
Zhao, Shiyu
Ma, Shiqing
Chen, Zhenyu
Computation and Language
Artificial Intelligence
Machine Learning
Reasoning is critical for large language models (LLMs) to excel in a wide range of tasks. While methods like Chain-of-Thought (CoT) reasoning and enhance LLM performance by decomposing problems into intermediate steps, they also incur significant overhead in token usage, leading to increased costs. We find that the reasoning process of current LLMs is unnecessarily lengthy and it can be compressed by including a reasonable token budget in the prompt, but the choice of token budget plays a crucial role in the actual compression effectiveness. We then propose a token-budget-aware LLM reasoning framework that dynamically adjusts the number of reasoning tokens based on the reasoning complexity of each problem. Experiments show that our method effectively reduces token costs in CoT reasoning with only a slight performance reduction, offering a practical solution to balance efficiency and accuracy in LLM reasoning. Code: https://github.com/GeniusHTX/TALE
title Token-Budget-Aware LLM Reasoning
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.18547