Active-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Muzhi, Zhong, Hao, Zhao, Canyu, Du, Zongze, Huang, Zheng, Liu, Mingyu, Chen, Hao, Zou, Cheng, Chen, Jingdong, Yang, Ming, Shen, Chunhua |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
by: Zhong, Hao, et al.
Published: (2025)
by: Zhong, Hao, et al.
Published: (2025)
Bridge Thinking and Acting: Unleashing Physical Potential of VLM with Generalizable Action Expert
by: Liu, Mingyu, et al.
Published: (2025)
by: Liu, Mingyu, et al.
Published: (2025)
NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
by: Huang, Zheng, et al.
Published: (2025)
by: Huang, Zheng, et al.
Published: (2025)
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
by: Zhao, Canyu, et al.
Published: (2025)
by: Zhao, Canyu, et al.
Published: (2025)
Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?
by: Li, Liyang, et al.
Published: (2026)
by: Li, Liyang, et al.
Published: (2026)
Generative Active Learning for Long-tailed Instance Segmentation
by: Zhu, Muzhi, et al.
Published: (2024)
by: Zhu, Muzhi, et al.
Published: (2024)
HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
by: Chen, Cong, et al.
Published: (2025)
by: Chen, Cong, et al.
Published: (2025)
Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality
by: Luo, Zekai, et al.
Published: (2025)
by: Luo, Zekai, et al.
Published: (2025)
GUI-Shepherd: Reliable Process Reward and Verification for Long-Sequence GUI Tasks
by: Chen, Cong, et al.
Published: (2025)
by: Chen, Cong, et al.
Published: (2025)
MovieDreamer: Hierarchical Generation for Coherent Long Visual Sequence
by: Zhao, Canyu, et al.
Published: (2024)
by: Zhao, Canyu, et al.
Published: (2024)
From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
by: Zhu, Wenxin, et al.
Published: (2025)
by: Zhu, Wenxin, et al.
Published: (2025)
PRISM-XR: Empowering Privacy-Aware XR Collaboration with Multimodal Large Language Models
by: Chen, Jiangong, et al.
Published: (2026)
by: Chen, Jiangong, et al.
Published: (2026)
OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering
by: Jia, Yiduo, et al.
Published: (2026)
by: Jia, Yiduo, et al.
Published: (2026)
Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization
by: Zhao, Canyu, et al.
Published: (2025)
by: Zhao, Canyu, et al.
Published: (2025)
Matcher: Segment Anything with One Shot Using All-Purpose Feature Matching
by: Liu, Yang, et al.
Published: (2023)
by: Liu, Yang, et al.
Published: (2023)
MARBLE: Multi-Aspect Reward Balance for Diffusion RL
by: Zhao, Canyu, et al.
Published: (2026)
by: Zhao, Canyu, et al.
Published: (2026)
StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
by: Liu, Mingyu, et al.
Published: (2025)
by: Liu, Mingyu, et al.
Published: (2025)
Empowering Language Models with Active Inquiry for Deeper Understanding
by: Pang, Jing-Cheng, et al.
Published: (2024)
by: Pang, Jing-Cheng, et al.
Published: (2024)
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
by: Li, Liyang, et al.
Published: (2026)
by: Li, Liyang, et al.
Published: (2026)
SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories
by: Zhu, Muzhi, et al.
Published: (2025)
by: Zhu, Muzhi, et al.
Published: (2025)
ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models
by: Wang, Ziyue, et al.
Published: (2024)
by: Wang, Ziyue, et al.
Published: (2024)
DiverGen: Improving Instance Segmentation by Learning Wider Data Distribution with More Diverse Generative Data
by: Fan, Chengxiang, et al.
Published: (2024)
by: Fan, Chengxiang, et al.
Published: (2024)
A Simple Image Segmentation Framework via In-Context Examples
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept Composition
by: Ding, Ganggui, et al.
Published: (2024)
by: Ding, Ganggui, et al.
Published: (2024)
Can Large Language Models Identify Authorship?
by: Huang, Baixiang, et al.
Published: (2024)
by: Huang, Baixiang, et al.
Published: (2024)
What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?
by: Xu, Guangkai, et al.
Published: (2024)
by: Xu, Guangkai, et al.
Published: (2024)
Reviving The Classics: Active Reward Modeling in Large Language Model Alignment
by: Shen, Yunyi, et al.
Published: (2025)
by: Shen, Yunyi, et al.
Published: (2025)
Active Layer-Contrastive Decoding Reduces Hallucination in Large Language Model Generation
by: Zhang, Hongxiang, et al.
Published: (2025)
by: Zhang, Hongxiang, et al.
Published: (2025)
S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
by: Dai, Muzhi, et al.
Published: (2025)
by: Dai, Muzhi, et al.
Published: (2025)
PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training
by: Chen, Cong, et al.
Published: (2025)
by: Chen, Cong, et al.
Published: (2025)
Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation
by: Zhu, Muzhi, et al.
Published: (2024)
by: Zhu, Muzhi, et al.
Published: (2024)
Exploring Spatial Intelligence from a Generative Perspective
by: Zhu, Muzhi, et al.
Published: (2026)
by: Zhu, Muzhi, et al.
Published: (2026)
Intelligent Power Grid Design Review via Active Perception-Enabled Multimodal Large Language Models
by: Tan, Taoliang, et al.
Published: (2025)
by: Tan, Taoliang, et al.
Published: (2025)
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
by: Sun, Yanpeng, et al.
Published: (2024)
by: Sun, Yanpeng, et al.
Published: (2024)
Active Visual Perception: Opportunities and Challenges
by: Li, Yian, et al.
Published: (2025)
by: Li, Yian, et al.
Published: (2025)
LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation
by: Duan, Yinglin, et al.
Published: (2025)
by: Duan, Yinglin, et al.
Published: (2025)
TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models
by: Zhou, Wenhao, et al.
Published: (2025)
by: Zhou, Wenhao, et al.
Published: (2025)
Unified Open-World Segmentation with Multi-Modal Prompts
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Cramer-Rao Bound Optimization for Active RIS-Empowered ISAC Systems
by: Zhu, Qi, et al.
Published: (2023)
by: Zhu, Qi, et al.
Published: (2023)
Chain of Visual Perception: Harnessing Multimodal Large Language Models for Zero-shot Camouflaged Object Detection
by: Tang, Lv, et al.
Published: (2023)
by: Tang, Lv, et al.
Published: (2023)
Similar Items
-
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
by: Zhong, Hao, et al.
Published: (2025) -
Bridge Thinking and Acting: Unleashing Physical Potential of VLM with Generalizable Action Expert
by: Liu, Mingyu, et al.
Published: (2025) -
NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
by: Huang, Zheng, et al.
Published: (2025) -
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
by: Zhao, Canyu, et al.
Published: (2025) -
Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?
by: Li, Liyang, et al.
Published: (2026)