Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yan, Shilin, Tong, Jintao, Xue, Hongwei, Tang, Xiaojun, Wang, Yangyang, Shi, Kunyu, Zhang, Guannan, Li, Ruixuan, Zou, Yixiong
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913019516157952
author Yan, Shilin
Tong, Jintao
Xue, Hongwei
Tang, Xiaojun
Wang, Yangyang
Shi, Kunyu
Zhang, Guannan
Li, Ruixuan
Zou, Yixiong
author_facet Yan, Shilin
Tong, Jintao
Xue, Hongwei
Tang, Xiaojun
Wang, Yangyang
Shi, Kunyu
Zhang, Guannan
Li, Ruixuan
Zou, Yixiong
contents The advent of agentic multimodal models has empowered systems to actively interact with external environments. However, current agents suffer from a profound meta-cognitive deficit: they struggle to arbitrate between leveraging internal knowledge and querying external utilities. Consequently, they frequently fall prey to blind tool invocation, resorting to reflexive tool execution even when queries are resolvable from the raw visual context. This pathological behavior precipitates severe latency bottlenecks and injects extraneous noise that derails sound reasoning. Existing reinforcement learning protocols attempt to mitigate this via a scalarized reward that penalizes tool usage. Yet, this coupled formulation creates an irreconcilable optimization dilemma: an aggressive penalty suppresses essential tool use, whereas a mild penalty is entirely subsumed by the variance of the accuracy reward during advantage normalization, rendering it impotent against tool overuse. To transcend this bottleneck, we propose HDPO, a framework that reframes tool efficiency from a competing scalar objective to a strictly conditional one. By eschewing reward scalarization, HDPO maintains two orthogonal optimization channels: an accuracy channel that maximizes task correctness, and an efficiency channel that enforces execution economy exclusively within accurate trajectories via conditional advantage estimation. This decoupled architecture naturally induces a cognitive curriculum-compelling the agent to first master task resolution before refining its self-reliance. Extensive evaluations demonstrate that our resulting model, Metis, reduces tool invocations by orders of magnitude while simultaneously elevating reasoning accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2604_08545
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models
Yan, Shilin
Tong, Jintao
Xue, Hongwei
Tang, Xiaojun
Wang, Yangyang
Shi, Kunyu
Zhang, Guannan
Li, Ruixuan
Zou, Yixiong
Computer Vision and Pattern Recognition
Artificial Intelligence
The advent of agentic multimodal models has empowered systems to actively interact with external environments. However, current agents suffer from a profound meta-cognitive deficit: they struggle to arbitrate between leveraging internal knowledge and querying external utilities. Consequently, they frequently fall prey to blind tool invocation, resorting to reflexive tool execution even when queries are resolvable from the raw visual context. This pathological behavior precipitates severe latency bottlenecks and injects extraneous noise that derails sound reasoning. Existing reinforcement learning protocols attempt to mitigate this via a scalarized reward that penalizes tool usage. Yet, this coupled formulation creates an irreconcilable optimization dilemma: an aggressive penalty suppresses essential tool use, whereas a mild penalty is entirely subsumed by the variance of the accuracy reward during advantage normalization, rendering it impotent against tool overuse. To transcend this bottleneck, we propose HDPO, a framework that reframes tool efficiency from a competing scalar objective to a strictly conditional one. By eschewing reward scalarization, HDPO maintains two orthogonal optimization channels: an accuracy channel that maximizes task correctness, and an efficiency channel that enforces execution economy exclusively within accurate trajectories via conditional advantage estimation. This decoupled architecture naturally induces a cognitive curriculum-compelling the agent to first master task resolution before refining its self-reliance. Extensive evaluations demonstrate that our resulting model, Metis, reduces tool invocations by orders of magnitude while simultaneously elevating reasoning accuracy.
title Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2604.08545