Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Dong, Xuanzhao, Zhu, Wenhui, Qiu, Peijie, Chen, Xiwen, Yu, Xiaobing, Li, Xin, Wang, Zhipeng, Tang, Shao, Li, Gen, Xiong, Yujian, Wang, Hao, Chen, Yanxi, Tiwari, Prayag, Wang, Yalin
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916053242609664
author Dong, Xuanzhao
Zhu, Wenhui
Qiu, Peijie
Chen, Xiwen
Yu, Xiaobing
Li, Xin
Wang, Zhipeng
Tang, Shao
Li, Gen
Xiong, Yujian
Wang, Hao
Chen, Yanxi
Tiwari, Prayag
Wang, Yalin
author_facet Dong, Xuanzhao
Zhu, Wenhui
Qiu, Peijie
Chen, Xiwen
Yu, Xiaobing
Li, Xin
Wang, Zhipeng
Tang, Shao
Li, Gen
Xiong, Yujian
Wang, Hao
Chen, Yanxi
Tiwari, Prayag
Wang, Yalin
contents Despite their popularity and success, Multimodal Large Language Models (MLLMs) often struggle to interpret images accurately, which limits their reasoning capability in complex scenarios (e.g., high object density and complex background clutter). Prior work mainly addresses this limitation by incorporating explicit visual cues like bounding boxes that require extra annotations. In addition, the resulting low-resolution crops often miss fine-grained details that MLLMs require for accurate reasoning. Therefore, we propose Mags-RL, an Agentic Reinforcement Learning (RL) framework that equips MLLMs with an external super-resolution "magnifying glass" agent for high-resolution fine-grained inspection. Specifically, the model performs two-round reasoning: in the first round, it generates an initial rationale and autonomously identifies regions of interest without relying on additional annotations; in the second round, it invokes a super-resolution agent to crop and upscale those regions, then revisits and verifies its earlier reasoning to produce the final answer. We also introduce a novel curriculum learning strategy that enables data-efficient RL training, needing as few as only 40 training samples to achieve reasonable performance. Experiments on VSR, TallyQA, and GQA subsets show its superior performance against recent strong competing methods, demonstrating high-quality reasoning with precise visual grounding. Code and weights will be released soon.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27960
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning
Dong, Xuanzhao
Zhu, Wenhui
Qiu, Peijie
Chen, Xiwen
Yu, Xiaobing
Li, Xin
Wang, Zhipeng
Tang, Shao
Li, Gen
Xiong, Yujian
Wang, Hao
Chen, Yanxi
Tiwari, Prayag
Wang, Yalin
Computer Vision and Pattern Recognition
Despite their popularity and success, Multimodal Large Language Models (MLLMs) often struggle to interpret images accurately, which limits their reasoning capability in complex scenarios (e.g., high object density and complex background clutter). Prior work mainly addresses this limitation by incorporating explicit visual cues like bounding boxes that require extra annotations. In addition, the resulting low-resolution crops often miss fine-grained details that MLLMs require for accurate reasoning. Therefore, we propose Mags-RL, an Agentic Reinforcement Learning (RL) framework that equips MLLMs with an external super-resolution "magnifying glass" agent for high-resolution fine-grained inspection. Specifically, the model performs two-round reasoning: in the first round, it generates an initial rationale and autonomously identifies regions of interest without relying on additional annotations; in the second round, it invokes a super-resolution agent to crop and upscale those regions, then revisits and verifies its earlier reasoning to produce the final answer. We also introduce a novel curriculum learning strategy that enables data-efficient RL training, needing as few as only 40 training samples to achieve reasonable performance. Experiments on VSR, TallyQA, and GQA subsets show its superior performance against recent strong competing methods, demonstrating high-quality reasoning with precise visual grounding. Code and weights will be released soon.
title Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.27960