AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Song, Mingyang, Sun, Haoyu, Gu, Jiawei, Li, Linjie, Xu, Luxin, Krishna, Ranjay, Cheng, Yu
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910002916098048
author Song, Mingyang
Sun, Haoyu
Gu, Jiawei
Li, Linjie
Xu, Luxin
Krishna, Ranjay
Cheng, Yu
author_facet Song, Mingyang
Sun, Haoyu
Gu, Jiawei
Li, Linjie
Xu, Luxin
Krishna, Ranjay
Cheng, Yu
contents When humans face problems beyond their immediate capabilities, they rely on tools, providing a promising paradigm for improving visual reasoning in multimodal large language models (MLLMs). Effective reasoning, therefore, hinges on knowing which tools to use, when to invoke them, and how to compose them over multiple steps, even when faced with new tools or new tasks. We introduce \textbf{AdaReasoner}, a family of multimodal models that learn tool use as a general reasoning skill rather than as tool-specific or explicitly supervised behavior. AdaReasoner is enabled by (i) a scalable data curation pipeline exposing models to long-horizon, multi-step tool interactions; (ii) Tool-GRPO, a reinforcement learning algorithm that optimizes tool selection and sequencing based on end-task success; and (iii) an adaptive learning mechanism that dynamically regulates tool usage. Together, these components allow models to infer tool utility from task context and intermediate outcomes, enabling coordination of multiple tools and generalization to unseen tools. Empirically, AdaReasoner exhibits strong tool-adaptive and generalization behaviors: it autonomously adopts beneficial tools, suppresses irrelevant ones, and adjusts tool usage frequency based on task demands, despite never being explicitly trained to do so. These capabilities translate into state-of-the-art performance across challenging benchmarks, improving the 7B base model by +24.9\% on average and surpassing strong proprietary systems such as GPT-5 on multiple tasks, including VSP and Jigsaw.
format Preprint
id arxiv_https___arxiv_org_abs_2601_18631
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning
Song, Mingyang
Sun, Haoyu
Gu, Jiawei
Li, Linjie
Xu, Luxin
Krishna, Ranjay
Cheng, Yu
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Multiagent Systems
When humans face problems beyond their immediate capabilities, they rely on tools, providing a promising paradigm for improving visual reasoning in multimodal large language models (MLLMs). Effective reasoning, therefore, hinges on knowing which tools to use, when to invoke them, and how to compose them over multiple steps, even when faced with new tools or new tasks. We introduce \textbf{AdaReasoner}, a family of multimodal models that learn tool use as a general reasoning skill rather than as tool-specific or explicitly supervised behavior. AdaReasoner is enabled by (i) a scalable data curation pipeline exposing models to long-horizon, multi-step tool interactions; (ii) Tool-GRPO, a reinforcement learning algorithm that optimizes tool selection and sequencing based on end-task success; and (iii) an adaptive learning mechanism that dynamically regulates tool usage. Together, these components allow models to infer tool utility from task context and intermediate outcomes, enabling coordination of multiple tools and generalization to unseen tools. Empirically, AdaReasoner exhibits strong tool-adaptive and generalization behaviors: it autonomously adopts beneficial tools, suppresses irrelevant ones, and adjusts tool usage frequency based on task demands, despite never being explicitly trained to do so. These capabilities translate into state-of-the-art performance across challenging benchmarks, improving the 7B base model by +24.9\% on average and surpassing strong proprietary systems such as GPT-5 on multiple tasks, including VSP and Jigsaw.
title AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Multiagent Systems
url https://arxiv.org/abs/2601.18631