One RL to See Them All: Visual Triple Unified Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Yan, Du, Linge, Shen, Xuyang, Chen, Shaoxiang, Li, Pengfei, Ren, Qibing, Ma, Lizhuang, Dai, Yuchao, Liu, Pengfei, Yan, Junjie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What Does Vision Tool-Use Reinforcement Learning Really Learn? Disentangling Tool-Induced and Intrinsic Effects for Crop-and-Zoom
by: Ma, Yan, et al.
Published: (2026)
by: Ma, Yan, et al.
Published: (2026)
Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme
by: Ma, Yan, et al.
Published: (2025)
by: Ma, Yan, et al.
Published: (2025)
One Sample to Rule Them All: Extreme Data Efficiency in Multidiscipline Reasoning with Reinforcement Learning
by: Li, Yiyuan, et al.
Published: (2026)
by: Li, Yiyuan, et al.
Published: (2026)
One Framework to Rule Them All: Unifying RL-Based and RL-Free Methods in RLHF
by: Cai, Xin
Published: (2025)
by: Cai, Xin
Published: (2025)
When AI Agents Collude Online: Financial Fraud Risks by Collaborative LLM Agents on Social Platforms
by: Ren, Qibing, et al.
Published: (2025)
by: Ren, Qibing, et al.
Published: (2025)
When Autonomy Goes Rogue: Preparing for Risks of Multi-Agent Collusion in Social Systems
by: Ren, Qibing, et al.
Published: (2025)
by: Ren, Qibing, et al.
Published: (2025)
One Ring to Rule Them All: Unifying Group-Based RL via Dynamic Power-Mean Geometry
by: Zhao, Weisong, et al.
Published: (2026)
by: Zhao, Weisong, et al.
Published: (2026)
CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion
by: Ren, Qibing, et al.
Published: (2024)
by: Ren, Qibing, et al.
Published: (2024)
OmniCaptioner: One Captioner to Rule Them All
by: Lu, Yiting, et al.
Published: (2025)
by: Lu, Yiting, et al.
Published: (2025)
TabPFN: One Model to Rule Them All?
by: Zhang, Qiong, et al.
Published: (2025)
by: Zhang, Qiong, et al.
Published: (2025)
MoPS: Modular Story Premise Synthesis for Open-Ended Automatic Story Generation
by: Ma, Yan, et al.
Published: (2024)
by: Ma, Yan, et al.
Published: (2024)
From Physical Degradation Models to Task-Aware All-in-One Image Restoration
by: Gao, Hu, et al.
Published: (2026)
by: Gao, Hu, et al.
Published: (2026)
Weak-to-Strong Reasoning
by: Yang, Yuqing, et al.
Published: (2024)
by: Yang, Yuqing, et al.
Published: (2024)
ToRL: Scaling Tool-Integrated RL
by: Li, Xuefeng, et al.
Published: (2025)
by: Li, Xuefeng, et al.
Published: (2025)
One Dinomaly2 Detect Them All: A Unified Framework for Full-Spectrum Unsupervised Anomaly Detection
by: Guo, Jia, et al.
Published: (2025)
by: Guo, Jia, et al.
Published: (2025)
LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts
by: Ren, Qibing, et al.
Published: (2024)
by: Ren, Qibing, et al.
Published: (2024)
One Whisper to Grade Them All
by: Phan, Nhan, et al.
Published: (2025)
by: Phan, Nhan, et al.
Published: (2025)
One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework
by: Bianchi, Lorenzo, et al.
Published: (2025)
by: Bianchi, Lorenzo, et al.
Published: (2025)
ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
by: Chern, Ethan, et al.
Published: (2024)
by: Chern, Ethan, et al.
Published: (2024)
One Diffusion to Generate Them All
by: Le, Duong H., et al.
Published: (2024)
by: Le, Duong H., et al.
Published: (2024)
One Framework to Rule Them All: Unifying Multimodal Tasks with LLM Neural-Tuning
by: Sun, Hao, et al.
Published: (2024)
by: Sun, Hao, et al.
Published: (2024)
ATATA: One Algorithm to Align Them All
by: Pang, Boyi, et al.
Published: (2026)
by: Pang, Boyi, et al.
Published: (2026)
StyleRWKV: High-Quality and High-Efficiency Style Transfer with RWKV-like Architecture
by: Dai, Miaomiao, et al.
Published: (2024)
by: Dai, Miaomiao, et al.
Published: (2024)
One Noise to Rule Them All: Learning a Unified Model of Spatially-Varying Noise Patterns
by: Maesumi, Arman, et al.
Published: (2024)
by: Maesumi, Arman, et al.
Published: (2024)
One Channel to Rule Them All: Rethinking Input Representation for Visual Place Recognition
by: Ismagilov, Timur, et al.
Published: (2026)
by: Ismagilov, Timur, et al.
Published: (2026)
Embed-RL: Reinforcement Learning for Reasoning-Driven Multimodal Embeddings
by: Jiang, Haonan, et al.
Published: (2026)
by: Jiang, Haonan, et al.
Published: (2026)
One Refiner to Unlock Them All: Inference-Time Reasoning Elicitation via Reinforcement Query Refinement
by: Zhou, Yixiao, et al.
Published: (2026)
by: Zhou, Yixiao, et al.
Published: (2026)
MIRG-RL: Multi-Image Reasoning and Grounding with Reinforcement Learning
by: Zheng, Lihao, et al.
Published: (2025)
by: Zheng, Lihao, et al.
Published: (2025)
MegaLoc: One Retrieval to Place Them All
by: Berton, Gabriele, et al.
Published: (2025)
by: Berton, Gabriele, et al.
Published: (2025)
SupeRANSAC: One RANSAC to Rule Them All
by: Barath, Daniel
Published: (2025)
by: Barath, Daniel
Published: (2025)
One LLM to Train Them All: Multi-Task Learning Framework for Fact-Checking
by: Larsson, Malin Astrid, et al.
Published: (2026)
by: Larsson, Malin Astrid, et al.
Published: (2026)
See then Tell: Enhancing Key Information Extraction with Vision Grounding
by: Liu, Shuhang, et al.
Published: (2024)
by: Liu, Shuhang, et al.
Published: (2024)
UVL2: A Unified Framework for Video Tampering Localization
by: Pei, Pengfei
Published: (2023)
by: Pei, Pengfei
Published: (2023)
Unleashing Degradation-Carrying Features in Symmetric U-Net: Simpler and Stronger Baselines for All-in-One Image Restoration
by: Jiao, Wenlong, et al.
Published: (2025)
by: Jiao, Wenlong, et al.
Published: (2025)
One Prompt To Rule Them All: LLMs for Opinion Summary Evaluation
by: Siledar, Tejpalsingh, et al.
Published: (2024)
by: Siledar, Tejpalsingh, et al.
Published: (2024)
UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
by: Jiao, Yang, et al.
Published: (2025)
by: Jiao, Yang, et al.
Published: (2025)
Driving Intents Amplify Planning-Oriented Reinforcement Learning
by: Lu, Hengtong, et al.
Published: (2026)
by: Lu, Hengtong, et al.
Published: (2026)
A Hyperdimensional One Place Signature to Represent Them All: Stackable Descriptors For Visual Place Recognition
by: Malone, Connor, et al.
Published: (2024)
by: Malone, Connor, et al.
Published: (2024)
You Only Scan Once: Efficient Multi-dimension Sequential Modeling with LightNet
by: Qin, Zhen, et al.
Published: (2024)
by: Qin, Zhen, et al.
Published: (2024)
LIMR: Less is More for RL Scaling
by: Li, Xuefeng, et al.
Published: (2025)
by: Li, Xuefeng, et al.
Published: (2025)
Similar Items
-
What Does Vision Tool-Use Reinforcement Learning Really Learn? Disentangling Tool-Induced and Intrinsic Effects for Crop-and-Zoom
by: Ma, Yan, et al.
Published: (2026) -
Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme
by: Ma, Yan, et al.
Published: (2025) -
One Sample to Rule Them All: Extreme Data Efficiency in Multidiscipline Reasoning with Reinforcement Learning
by: Li, Yiyuan, et al.
Published: (2026) -
One Framework to Rule Them All: Unifying RL-Based and RL-Free Methods in RLHF
by: Cai, Xin
Published: (2025) -
When AI Agents Collude Online: Financial Fraud Risks by Collaborative LLM Agents on Social Platforms
by: Ren, Qibing, et al.
Published: (2025)