ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Shengyuan, Fang, Xinyu, Liu, Ziyu, Zang, Yuhang, Cao, Yuhang, Zhao, Xiangyu, Duan, Haodong, Dong, Xiaoyi, Liang, Jianze, Wang, Bin, He, Conghui, Lin, Dahua, Wang, Jiaqi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915654347522048
author Ding, Shengyuan
Fang, Xinyu
Liu, Ziyu
Zang, Yuhang
Cao, Yuhang
Zhao, Xiangyu
Duan, Haodong
Dong, Xiaoyi
Liang, Jianze
Wang, Bin
He, Conghui
Lin, Dahua
Wang, Jiaqi
author_facet Ding, Shengyuan
Fang, Xinyu
Liu, Ziyu
Zang, Yuhang
Cao, Yuhang
Zhao, Xiangyu
Duan, Haodong
Dong, Xiaoyi
Liang, Jianze
Wang, Bin
He, Conghui
Lin, Dahua
Wang, Jiaqi
contents Reward models are critical for aligning vision-language systems with human preferences, yet current approaches suffer from hallucination, weak visual grounding, and an inability to use tools for verification, limiting their reliability on complex multimodal reasoning tasks. We present ARM-Thinker, an A}gentic multimodal Reward Model that autonomously invokes external tools (e.g., image cropping, doc page retrieval) to ground judgments in verifiable evidence, replacing static, non-interactive reward scoring. This enables the model to verify fine-grained visual details, cross-reference multi-page evidence, and validate reasoning claims, which are capabilities absent in existing reward models. We train ARM-Thinker with multi-stage reinforcement learning, jointly optimizing tool-calling decisions and judgment accuracy. To evaluate agentic reward modeling, we introduce ARMBench-VL, comprising three benchmarks that assess fine-grained visual grounding (image-level tools), multi-page document understanding (retrieval tools), and instruction following (text-level verification). ARM-Thinker achieves +16.2% average improvement on reward modeling benchmarks, +9.6% on tool-use tasks, and outperforms baselines on multimodal math and logical reasoning benchmarks. Our results demonstrate that agentic capabilities significantly enhance both accuracy and interpretability of reward models.
format Preprint
id arxiv_https___arxiv_org_abs_2512_05111
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
Ding, Shengyuan
Fang, Xinyu
Liu, Ziyu
Zang, Yuhang
Cao, Yuhang
Zhao, Xiangyu
Duan, Haodong
Dong, Xiaoyi
Liang, Jianze
Wang, Bin
He, Conghui
Lin, Dahua
Wang, Jiaqi
Computer Vision and Pattern Recognition
Reward models are critical for aligning vision-language systems with human preferences, yet current approaches suffer from hallucination, weak visual grounding, and an inability to use tools for verification, limiting their reliability on complex multimodal reasoning tasks. We present ARM-Thinker, an A}gentic multimodal Reward Model that autonomously invokes external tools (e.g., image cropping, doc page retrieval) to ground judgments in verifiable evidence, replacing static, non-interactive reward scoring. This enables the model to verify fine-grained visual details, cross-reference multi-page evidence, and validate reasoning claims, which are capabilities absent in existing reward models. We train ARM-Thinker with multi-stage reinforcement learning, jointly optimizing tool-calling decisions and judgment accuracy. To evaluate agentic reward modeling, we introduce ARMBench-VL, comprising three benchmarks that assess fine-grained visual grounding (image-level tools), multi-page document understanding (retrieval tools), and instruction following (text-level verification). ARM-Thinker achieves +16.2% average improvement on reward modeling benchmarks, +9.6% on tool-use tasks, and outperforms baselines on multimodal math and logical reasoning benchmarks. Our results demonstrate that agentic capabilities significantly enhance both accuracy and interpretability of reward models.
title ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.05111