Active-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Muzhi, Zhong, Hao, Zhao, Canyu, Du, Zongze, Huang, Zheng, Liu, Mingyu, Chen, Hao, Zou, Cheng, Chen, Jingdong, Yang, Ming, Shen, Chunhua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916762036994048
author Zhu, Muzhi
Zhong, Hao
Zhao, Canyu
Du, Zongze
Huang, Zheng
Liu, Mingyu
Chen, Hao
Zou, Cheng
Chen, Jingdong
Yang, Ming
Shen, Chunhua
author_facet Zhu, Muzhi
Zhong, Hao
Zhao, Canyu
Du, Zongze
Huang, Zheng
Liu, Mingyu
Chen, Hao
Zou, Cheng
Chen, Jingdong
Yang, Ming
Shen, Chunhua
contents Active vision, also known as active perception, refers to the process of actively selecting where and how to look in order to gather task-relevant information. It is a critical component of efficient perception and decision-making in humans and advanced embodied agents. Recently, the use of Multimodal Large Language Models (MLLMs) as central planning and decision-making modules in robotic systems has gained extensive attention. However, despite the importance of active perception in embodied intelligence, there is little to no exploration of how MLLMs can be equipped with or learn active perception capabilities. In this paper, we first provide a systematic definition of MLLM-based active perception tasks. We point out that the recently proposed GPT-o3 model's zoom-in search strategy can be regarded as a special case of active perception; however, it still suffers from low search efficiency and inaccurate region selection. To address these issues, we propose ACTIVE-O3, a purely reinforcement learning based training framework built on top of GRPO, designed to equip MLLMs with active perception capabilities. We further establish a comprehensive benchmark suite to evaluate ACTIVE-O3 across both general open-world tasks, such as small-object and dense object grounding, and domain-specific scenarios, including small object detection in remote sensing and autonomous driving, as well as fine-grained interactive segmentation. In addition, ACTIVE-O3 also demonstrates strong zero-shot reasoning abilities on the V* Benchmark, without relying on any explicit reasoning data. We hope that our work can provide a simple codebase and evaluation protocol to facilitate future research on active perception in MLLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21457
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Active-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO
Zhu, Muzhi
Zhong, Hao
Zhao, Canyu
Du, Zongze
Huang, Zheng
Liu, Mingyu
Chen, Hao
Zou, Cheng
Chen, Jingdong
Yang, Ming
Shen, Chunhua
Computer Vision and Pattern Recognition
Artificial Intelligence
Active vision, also known as active perception, refers to the process of actively selecting where and how to look in order to gather task-relevant information. It is a critical component of efficient perception and decision-making in humans and advanced embodied agents. Recently, the use of Multimodal Large Language Models (MLLMs) as central planning and decision-making modules in robotic systems has gained extensive attention. However, despite the importance of active perception in embodied intelligence, there is little to no exploration of how MLLMs can be equipped with or learn active perception capabilities. In this paper, we first provide a systematic definition of MLLM-based active perception tasks. We point out that the recently proposed GPT-o3 model's zoom-in search strategy can be regarded as a special case of active perception; however, it still suffers from low search efficiency and inaccurate region selection. To address these issues, we propose ACTIVE-O3, a purely reinforcement learning based training framework built on top of GRPO, designed to equip MLLMs with active perception capabilities. We further establish a comprehensive benchmark suite to evaluate ACTIVE-O3 across both general open-world tasks, such as small-object and dense object grounding, and domain-specific scenarios, including small object detection in remote sensing and autonomous driving, as well as fine-grained interactive segmentation. In addition, ACTIVE-O3 also demonstrates strong zero-shot reasoning abilities on the V* Benchmark, without relying on any explicit reasoning data. We hope that our work can provide a simple codebase and evaluation protocol to facilitate future research on active perception in MLLMs.
title Active-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.21457