PyVision-RL: Forging Open Agentic Vision Models via RL

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhao, Shitian, Lin, Shaoheng, Li, Ming, Zhang, Haoquan, Peng, Wenshuo, Zhang, Kaipeng, Wei, Chen
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912923157266432
author Zhao, Shitian
Lin, Shaoheng
Li, Ming
Zhang, Haoquan
Peng, Wenshuo
Zhang, Kaipeng
Wei, Chen
author_facet Zhao, Shitian
Lin, Shaoheng
Li, Ming
Zhang, Haoquan
Peng, Wenshuo
Zhang, Kaipeng
Wei, Chen
contents Reinforcement learning for agentic multimodal models often suffers from interaction collapse, where models learn to reduce tool usage and multi-turn reasoning, limiting the benefits of agentic behavior. We introduce PyVision-RL, a reinforcement learning framework for open-weight multimodal models that stabilizes training and sustains interaction. Our approach combines an oversampling-filtering-ranking rollout strategy with an accumulative tool reward to prevent collapse and encourage multi-turn tool use. Using a unified training pipeline, we develop PyVision-Image and PyVision-Video for image and video understanding. For video reasoning, PyVision-Video employs on-demand context construction, selectively sampling task-relevant frames during reasoning to significantly reduce visual token usage. Experiments show strong performance and improved efficiency, demonstrating that sustained interaction and on-demand visual processing are critical for scalable multimodal agents.
format Preprint
id arxiv_https___arxiv_org_abs_2602_20739
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PyVision-RL: Forging Open Agentic Vision Models via RL
Zhao, Shitian
Lin, Shaoheng
Li, Ming
Zhang, Haoquan
Peng, Wenshuo
Zhang, Kaipeng
Wei, Chen
Artificial Intelligence
Computer Vision and Pattern Recognition
Reinforcement learning for agentic multimodal models often suffers from interaction collapse, where models learn to reduce tool usage and multi-turn reasoning, limiting the benefits of agentic behavior. We introduce PyVision-RL, a reinforcement learning framework for open-weight multimodal models that stabilizes training and sustains interaction. Our approach combines an oversampling-filtering-ranking rollout strategy with an accumulative tool reward to prevent collapse and encourage multi-turn tool use. Using a unified training pipeline, we develop PyVision-Image and PyVision-Video for image and video understanding. For video reasoning, PyVision-Video employs on-demand context construction, selectively sampling task-relevant frames during reasoning to significantly reduce visual token usage. Experiments show strong performance and improved efficiency, demonstrating that sustained interaction and on-demand visual processing are critical for scalable multimodal agents.
title PyVision-RL: Forging Open Agentic Vision Models via RL
topic Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.20739