Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Tan, Reuben, Peng, Baolin, Yang, Zhengyuan, Cheng, Hao, Mees, Oier, Zhao, Theodore, Tupini, Andrea, Meijier, Isar, Wu, Qianhui, Yang, Yuncong, Liden, Lars, Gu, Yu, Zhang, Sheng, Liu, Xiaodong, Wang, Lijuan, Pollefeys, Marc, Lee, Yong Jae, Gao, Jianfeng
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911602342625280
author Tan, Reuben
Peng, Baolin
Yang, Zhengyuan
Cheng, Hao
Mees, Oier
Zhao, Theodore
Tupini, Andrea
Meijier, Isar
Wu, Qianhui
Yang, Yuncong
Liden, Lars
Gu, Yu
Zhang, Sheng
Liu, Xiaodong
Wang, Lijuan
Pollefeys, Marc
Lee, Yong Jae
Gao, Jianfeng
author_facet Tan, Reuben
Peng, Baolin
Yang, Zhengyuan
Cheng, Hao
Mees, Oier
Zhao, Theodore
Tupini, Andrea
Meijier, Isar
Wu, Qianhui
Yang, Yuncong
Liden, Lars
Gu, Yu
Zhang, Sheng
Liu, Xiaodong
Wang, Lijuan
Pollefeys, Marc
Lee, Yong Jae
Gao, Jianfeng
contents Agentic reasoning models trained with multimodal reinforcement learning (MMRL) have become increasingly capable, yet they are almost universally optimized using sparse, outcome-based rewards computed based on the final answers. Richer rewards computed from the reasoning tokens can improve learning significantly by providing more fine-grained guidance. However, it is challenging to compute more informative rewards in MMRL beyond those based on outcomes since different samples may require different scoring functions and teacher models may provide noisy reward signals too. In this paper, we introduce the Argos (Agentic Reward for Grounded & Objective Scoring), a principled reward agent to train multimodal reasoning models for agentic tasks. For each sample, Argos selects from a pool of teacher-model derived and rule-based scoring functions to simultaneously evaluate: (i) final response accuracy, (ii) spatiotemporal localization of referred entities and actions, and (iii) the quality of the reasoning process. We find that by leveraging our agentic verifier across both SFT data curation and RL training, our model achieves state-of-the-art results across multiple agentic tasks such as spatial reasoning, visual hallucination as well as robotics and embodied AI benchmarks. Critically, we demonstrate that just relying on SFT post-training on highly curated reasoning data is insufficient, as agents invariably collapse to ungrounded solutions during RL without our online verification. We also show that our agentic verifier can help to reduce reward-hacking in MMRL. Finally, we also provide a theoretical justification for the effectiveness of Argos through the concept of pareto-optimality.
format Preprint
id arxiv_https___arxiv_org_abs_2512_03438
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents
Tan, Reuben
Peng, Baolin
Yang, Zhengyuan
Cheng, Hao
Mees, Oier
Zhao, Theodore
Tupini, Andrea
Meijier, Isar
Wu, Qianhui
Yang, Yuncong
Liden, Lars
Gu, Yu
Zhang, Sheng
Liu, Xiaodong
Wang, Lijuan
Pollefeys, Marc
Lee, Yong Jae
Gao, Jianfeng
Artificial Intelligence
Agentic reasoning models trained with multimodal reinforcement learning (MMRL) have become increasingly capable, yet they are almost universally optimized using sparse, outcome-based rewards computed based on the final answers. Richer rewards computed from the reasoning tokens can improve learning significantly by providing more fine-grained guidance. However, it is challenging to compute more informative rewards in MMRL beyond those based on outcomes since different samples may require different scoring functions and teacher models may provide noisy reward signals too. In this paper, we introduce the Argos (Agentic Reward for Grounded & Objective Scoring), a principled reward agent to train multimodal reasoning models for agentic tasks. For each sample, Argos selects from a pool of teacher-model derived and rule-based scoring functions to simultaneously evaluate: (i) final response accuracy, (ii) spatiotemporal localization of referred entities and actions, and (iii) the quality of the reasoning process. We find that by leveraging our agentic verifier across both SFT data curation and RL training, our model achieves state-of-the-art results across multiple agentic tasks such as spatial reasoning, visual hallucination as well as robotics and embodied AI benchmarks. Critically, we demonstrate that just relying on SFT post-training on highly curated reasoning data is insufficient, as agents invariably collapse to ungrounded solutions during RL without our online verification. We also show that our agentic verifier can help to reduce reward-hacking in MMRL. Finally, we also provide a theoretical justification for the effectiveness of Argos through the concept of pareto-optimality.
title Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents
topic Artificial Intelligence
url https://arxiv.org/abs/2512.03438