One RL to See Them All: Visual Triple Unified Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Yan, Du, Linge, Shen, Xuyang, Chen, Shaoxiang, Li, Pengfei, Ren, Qibing, Ma, Lizhuang, Dai, Yuchao, Liu, Pengfei, Yan, Junjie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908968538865664
author Ma, Yan
Du, Linge
Shen, Xuyang
Chen, Shaoxiang
Li, Pengfei
Ren, Qibing
Ma, Lizhuang
Dai, Yuchao
Liu, Pengfei
Yan, Junjie
author_facet Ma, Yan
Du, Linge
Shen, Xuyang
Chen, Shaoxiang
Li, Pengfei
Ren, Qibing
Ma, Lizhuang
Dai, Yuchao
Liu, Pengfei
Yan, Junjie
contents Reinforcement learning (RL) is becoming an important direction for post-training vision-language models (VLMs), but public training methodologies for unified multimodal RL remain much less mature, especially for heterogeneous reasoning and perception-heavy tasks. We propose V-Triune, a Visual Triple Unified Reinforcement Learning methodology for unified multimodal RL. It organizes training around three coordinated abstractions: Sample-Level Reward Routing, Verifier-Level Outcome Verification, and Source-Level Diagnostics. Within this methodology, Dynamic IoU provides localization-specific reward shaping that avoids reward ambiguity under loose thresholds and reward sparsity under strict ones. Built on V-Triune, we develop Orsta (7B, 32B), a family of models jointly trained on eight reasoning and perception tasks. Under matched budgets, unified training matches or outperforms specialist mixtures. The final Orsta models improve over their backbones on MEGA-Bench, compare favorably with strong multi-task RL-VLM baselines, and transfer these gains to a broad set of downstream benchmarks. These results show that unified RL can improve both reasoning and perception within a single VLM RL pipeline.The V-Triune system, along with the Orsta models, is publicly available at https://github.com/MiniMax-AI/One-RL-to-See-Them-All.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18129
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle One RL to See Them All: Visual Triple Unified Reinforcement Learning
Ma, Yan
Du, Linge
Shen, Xuyang
Chen, Shaoxiang
Li, Pengfei
Ren, Qibing
Ma, Lizhuang
Dai, Yuchao
Liu, Pengfei
Yan, Junjie
Computer Vision and Pattern Recognition
Computation and Language
Reinforcement learning (RL) is becoming an important direction for post-training vision-language models (VLMs), but public training methodologies for unified multimodal RL remain much less mature, especially for heterogeneous reasoning and perception-heavy tasks. We propose V-Triune, a Visual Triple Unified Reinforcement Learning methodology for unified multimodal RL. It organizes training around three coordinated abstractions: Sample-Level Reward Routing, Verifier-Level Outcome Verification, and Source-Level Diagnostics. Within this methodology, Dynamic IoU provides localization-specific reward shaping that avoids reward ambiguity under loose thresholds and reward sparsity under strict ones. Built on V-Triune, we develop Orsta (7B, 32B), a family of models jointly trained on eight reasoning and perception tasks. Under matched budgets, unified training matches or outperforms specialist mixtures. The final Orsta models improve over their backbones on MEGA-Bench, compare favorably with strong multi-task RL-VLM baselines, and transfer these gains to a broad set of downstream benchmarks. These results show that unified RL can improve both reasoning and perception within a single VLM RL pipeline.The V-Triune system, along with the Orsta models, is publicly available at https://github.com/MiniMax-AI/One-RL-to-See-Them-All.
title One RL to See Them All: Visual Triple Unified Reinforcement Learning
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2505.18129