Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Jinghui, Yu, Haiyang, Xu, Siliang, Ran, Shiwei, Tang, Guozhi, Wang, Siqi, Shan, Bin, Fu, Teng, Feng, Hao, Tang, Jingqun, Wang, Han, Huang, Can
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916748301697024
author Lu, Jinghui
Yu, Haiyang
Xu, Siliang
Ran, Shiwei
Tang, Guozhi
Wang, Siqi
Shan, Bin
Fu, Teng
Feng, Hao
Tang, Jingqun
Wang, Han
Huang, Can
author_facet Lu, Jinghui
Yu, Haiyang
Xu, Siliang
Ran, Shiwei
Tang, Guozhi
Wang, Siqi
Shan, Bin
Fu, Teng
Feng, Hao
Tang, Jingqun
Wang, Han
Huang, Can
contents Recent advancements in reasoning have significantly enhanced the capabilities of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) across diverse tasks. However, excessive reliance on chain-of-thought (CoT) reasoning can impair model performance and brings unnecessarily lengthened outputs, reducing efficiency. Our work reveals that prolonged reasoning does not universally improve accuracy and even degrade performance on simpler tasks. To address this, we propose Certainty-based Adaptive Reasoning (CAR), a novel framework that dynamically switches between short answers and long-form reasoning based on the model perplexity. CAR first generates a short answer and evaluates its perplexity, triggering reasoning only when the model exhibits low confidence (i.e., high perplexity). Experiments across diverse multimodal VQA/KIE benchmarks and text reasoning datasets show that CAR outperforms both short-answer and long-form reasoning approaches, striking an optimal balance between accuracy and efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15154
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
Lu, Jinghui
Yu, Haiyang
Xu, Siliang
Ran, Shiwei
Tang, Guozhi
Wang, Siqi
Shan, Bin
Fu, Teng
Feng, Hao
Tang, Jingqun
Wang, Han
Huang, Can
Computation and Language
Artificial Intelligence
Multimedia
Recent advancements in reasoning have significantly enhanced the capabilities of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) across diverse tasks. However, excessive reliance on chain-of-thought (CoT) reasoning can impair model performance and brings unnecessarily lengthened outputs, reducing efficiency. Our work reveals that prolonged reasoning does not universally improve accuracy and even degrade performance on simpler tasks. To address this, we propose Certainty-based Adaptive Reasoning (CAR), a novel framework that dynamically switches between short answers and long-form reasoning based on the model perplexity. CAR first generates a short answer and evaluates its perplexity, triggering reasoning only when the model exhibits low confidence (i.e., high perplexity). Experiments across diverse multimodal VQA/KIE benchmarks and text reasoning datasets show that CAR outperforms both short-answer and long-form reasoning approaches, striking an optimal balance between accuracy and efficiency.
title Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
topic Computation and Language
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2505.15154