SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chng, Yong Xien, Hu, Tao, Tong, Wenwen, Li, Xueheng, Chen, Jiandong, Yu, Haojia, Lu, Jiefan, Guo, Hewei, Deng, Hanming, Xie, Chengjun, Huang, Gao, Lin, Dahua, Lu, Lewei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908784746561536
author Chng, Yong Xien
Hu, Tao
Tong, Wenwen
Li, Xueheng
Chen, Jiandong
Yu, Haojia
Lu, Jiefan
Guo, Hewei
Deng, Hanming
Xie, Chengjun
Huang, Gao
Lin, Dahua
Lu, Lewei
author_facet Chng, Yong Xien
Hu, Tao
Tong, Wenwen
Li, Xueheng
Chen, Jiandong
Yu, Haojia
Lu, Jiefan
Guo, Hewei
Deng, Hanming
Xie, Chengjun
Huang, Gao
Lin, Dahua
Lu, Lewei
contents While Vision-Language Models (VLMs) can solve complex tasks through agentic reasoning, their capabilities remain largely constrained to text-oriented chain-of-thought or isolated tool invocation. They fail to exhibit the human-like proficiency required to seamlessly interleave dynamic tool manipulation with continuous reasoning, particularly in knowledge-intensive and visually complex scenarios that demand coordinated external tools such as search and image cropping. In this work, we introduce SenseNova-MARS, a novel Multimodal Agentic Reasoning and Search framework that empowers VLMs with interleaved visual reasoning and tool-use capabilities via reinforcement learning (RL). Specifically, SenseNova-MARS dynamically integrates the image search, text search, and image crop tools to tackle fine-grained and knowledge-intensive visual understanding challenges. In the RL stage, we propose the Batch-Normalized Group Sequence Policy Optimization (BN-GSPO) algorithm to improve the training stability and advance the model's ability to invoke tools and reason effectively. To comprehensively evaluate the agentic VLMs on complex visual tasks, we introduce the HR-MMSearch benchmark, the first search-oriented benchmark composed of high-resolution images with knowledge-intensive and search-driven questions. Experiments demonstrate that SenseNova-MARS achieves state-of-the-art performance on open-source search and fine-grained image understanding benchmarks. Specifically, on search-oriented benchmarks, SenseNova-MARS-32B scores 74.3 on MMSearch and 54.4 on HR-MMSearch, surpassing proprietary models such as Gemini-3-Pro and GPT-5.2. SenseNova-MARS represents a promising step toward agentic VLMs by providing effective and robust tool-use capabilities. To facilitate further research in this field, we will release all code, models, and datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2512_24330
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning
Chng, Yong Xien
Hu, Tao
Tong, Wenwen
Li, Xueheng
Chen, Jiandong
Yu, Haojia
Lu, Jiefan
Guo, Hewei
Deng, Hanming
Xie, Chengjun
Huang, Gao
Lin, Dahua
Lu, Lewei
Computer Vision and Pattern Recognition
While Vision-Language Models (VLMs) can solve complex tasks through agentic reasoning, their capabilities remain largely constrained to text-oriented chain-of-thought or isolated tool invocation. They fail to exhibit the human-like proficiency required to seamlessly interleave dynamic tool manipulation with continuous reasoning, particularly in knowledge-intensive and visually complex scenarios that demand coordinated external tools such as search and image cropping. In this work, we introduce SenseNova-MARS, a novel Multimodal Agentic Reasoning and Search framework that empowers VLMs with interleaved visual reasoning and tool-use capabilities via reinforcement learning (RL). Specifically, SenseNova-MARS dynamically integrates the image search, text search, and image crop tools to tackle fine-grained and knowledge-intensive visual understanding challenges. In the RL stage, we propose the Batch-Normalized Group Sequence Policy Optimization (BN-GSPO) algorithm to improve the training stability and advance the model's ability to invoke tools and reason effectively. To comprehensively evaluate the agentic VLMs on complex visual tasks, we introduce the HR-MMSearch benchmark, the first search-oriented benchmark composed of high-resolution images with knowledge-intensive and search-driven questions. Experiments demonstrate that SenseNova-MARS achieves state-of-the-art performance on open-source search and fine-grained image understanding benchmarks. Specifically, on search-oriented benchmarks, SenseNova-MARS-32B scores 74.3 on MMSearch and 54.4 on HR-MMSearch, surpassing proprietary models such as Gemini-3-Pro and GPT-5.2. SenseNova-MARS represents a promising step toward agentic VLMs by providing effective and robust tool-use capabilities. To facilitate further research in this field, we will release all code, models, and datasets.
title SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.24330