MAXS: Meta-Adaptive Exploration with LLM Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jian, Wang, Zhiyuan, Wang, Zhangqi, He, Yu, Luo, Haoran, yuan, li, Zhang, Lingling, Mao, Rui, Lin, Qika, Liu, Jun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908763976368128
author Zhang, Jian
Wang, Zhiyuan
Wang, Zhangqi
He, Yu
Luo, Haoran
yuan, li
Zhang, Lingling
Mao, Rui
Lin, Qika
Liu, Jun
author_facet Zhang, Jian
Wang, Zhiyuan
Wang, Zhangqi
He, Yu
Luo, Haoran
yuan, li
Zhang, Lingling
Mao, Rui
Lin, Qika
Liu, Jun
contents Large Language Model (LLM) Agents exhibit inherent reasoning abilities through the collaboration of multiple tools. However, during agent inference, existing methods often suffer from (i) locally myopic generation, due to the absence of lookahead, and (ii) trajectory instability, where minor early errors can escalate into divergent reasoning paths. These issues make it difficult to balance global effectiveness and computational efficiency. To address these two issues, we propose meta-adaptive exploration with LLM agents https://github.com/exoskeletonzj/MAXS, a meta-adaptive reasoning framework based on LLM Agents that flexibly integrates tool execution and reasoning planning. MAXS employs a lookahead strategy to extend reasoning paths a few steps ahead, estimating the advantage value of tool usage, and combines step consistency variance and inter-step trend slopes to jointly select stable, consistent, and high-value reasoning steps. Additionally, we introduce a trajectory convergence mechanism that controls computational cost by halting further rollouts once path consistency is achieved, enabling a balance between resource efficiency and global effectiveness in multi-tool reasoning. We conduct extensive empirical studies across three base models (MiMo-VL-7B, Qwen2.5-VL-7B, Qwen2.5-VL-32B) and five datasets, demonstrating that MAXS consistently outperforms existing methods in both performance and inference efficiency. Further analysis confirms the effectiveness of our lookahead strategy and tool usage.
format Preprint
id arxiv_https___arxiv_org_abs_2601_09259
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MAXS: Meta-Adaptive Exploration with LLM Agents
Zhang, Jian
Wang, Zhiyuan
Wang, Zhangqi
He, Yu
Luo, Haoran
yuan, li
Zhang, Lingling
Mao, Rui
Lin, Qika
Liu, Jun
Artificial Intelligence
Large Language Model (LLM) Agents exhibit inherent reasoning abilities through the collaboration of multiple tools. However, during agent inference, existing methods often suffer from (i) locally myopic generation, due to the absence of lookahead, and (ii) trajectory instability, where minor early errors can escalate into divergent reasoning paths. These issues make it difficult to balance global effectiveness and computational efficiency. To address these two issues, we propose meta-adaptive exploration with LLM agents https://github.com/exoskeletonzj/MAXS, a meta-adaptive reasoning framework based on LLM Agents that flexibly integrates tool execution and reasoning planning. MAXS employs a lookahead strategy to extend reasoning paths a few steps ahead, estimating the advantage value of tool usage, and combines step consistency variance and inter-step trend slopes to jointly select stable, consistent, and high-value reasoning steps. Additionally, we introduce a trajectory convergence mechanism that controls computational cost by halting further rollouts once path consistency is achieved, enabling a balance between resource efficiency and global effectiveness in multi-tool reasoning. We conduct extensive empirical studies across three base models (MiMo-VL-7B, Qwen2.5-VL-7B, Qwen2.5-VL-32B) and five datasets, demonstrating that MAXS consistently outperforms existing methods in both performance and inference efficiency. Further analysis confirms the effectiveness of our lookahead strategy and tool usage.
title MAXS: Meta-Adaptive Exploration with LLM Agents
topic Artificial Intelligence
url https://arxiv.org/abs/2601.09259