AdaReasoner: Adaptive Reasoning Enables More Flexible Thinking in Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Xiangqi, Huang, Yue, Wang, Yanbo, Luo, Xiaonan, Guo, Kehan, Zhou, Yujun, Zhang, Xiangliang
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912640023920640
author Wang, Xiangqi
Huang, Yue
Wang, Yanbo
Luo, Xiaonan
Guo, Kehan
Zhou, Yujun
Zhang, Xiangliang
author_facet Wang, Xiangqi
Huang, Yue
Wang, Yanbo
Luo, Xiaonan
Guo, Kehan
Zhou, Yujun
Zhang, Xiangliang
contents LLMs often need effective configurations, like temperature and reasoning steps, to handle tasks requiring sophisticated reasoning and problem-solving, ranging from joke generation to mathematical reasoning. Existing prompting approaches usually adopt general-purpose, fixed configurations that work 'well enough' across tasks but seldom achieve task-specific optimality. To address this gap, we introduce AdaReasoner, an LLM-agnostic plugin designed for any LLM to automate adaptive reasoning configurations for tasks requiring different types of thinking. AdaReasoner is trained using a reinforcement learning (RL) framework, combining a factorized action space with a targeted exploration strategy, along with a pretrained reward model to optimize the policy model for reasoning configurations with only a few-shot guide. AdaReasoner is backed by theoretical guarantees and experiments of fast convergence and a sublinear policy gap. Across six different LLMs and a variety of reasoning tasks, it consistently outperforms standard baselines, preserves out-of-distribution robustness, and yield gains on knowledge-intensive tasks through tailored prompts.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17312
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AdaReasoner: Adaptive Reasoning Enables More Flexible Thinking in Large Language Models
Wang, Xiangqi
Huang, Yue
Wang, Yanbo
Luo, Xiaonan
Guo, Kehan
Zhou, Yujun
Zhang, Xiangliang
Artificial Intelligence
Machine Learning
LLMs often need effective configurations, like temperature and reasoning steps, to handle tasks requiring sophisticated reasoning and problem-solving, ranging from joke generation to mathematical reasoning. Existing prompting approaches usually adopt general-purpose, fixed configurations that work 'well enough' across tasks but seldom achieve task-specific optimality. To address this gap, we introduce AdaReasoner, an LLM-agnostic plugin designed for any LLM to automate adaptive reasoning configurations for tasks requiring different types of thinking. AdaReasoner is trained using a reinforcement learning (RL) framework, combining a factorized action space with a targeted exploration strategy, along with a pretrained reward model to optimize the policy model for reasoning configurations with only a few-shot guide. AdaReasoner is backed by theoretical guarantees and experiments of fast convergence and a sublinear policy gap. Across six different LLMs and a variety of reasoning tasks, it consistently outperforms standard baselines, preserves out-of-distribution robustness, and yield gains on knowledge-intensive tasks through tailored prompts.
title AdaReasoner: Adaptive Reasoning Enables More Flexible Thinking in Large Language Models
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.17312