SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shu, Fangxun, Ye, Yongjie, Liao, Yue, Kang, Zijian, Yin, Weijie, Wang, Jiacong, Liang, Xiao, Yan, Shuicheng, Feng, Chao
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915769378406400
author Shu, Fangxun
Ye, Yongjie
Liao, Yue
Kang, Zijian
Yin, Weijie
Wang, Jiacong
Liang, Xiao
Yan, Shuicheng
Feng, Chao
author_facet Shu, Fangxun
Ye, Yongjie
Liao, Yue
Kang, Zijian
Yin, Weijie
Wang, Jiacong
Liang, Xiao
Yan, Shuicheng
Feng, Chao
contents We introduce SAIL-RL, a reinforcement learning (RL) post-training framework that enhances the reasoning capabilities of multimodal large language models (MLLMs) by teaching them when and how to think. Existing approaches are limited by outcome-only supervision, which rewards correct answers without ensuring sound reasoning, and by uniform thinking strategies, which often lead to overthinking on simple tasks and underthinking on complex ones. SAIL-RL addresses these challenges with a dual reward system: the Thinking Reward, which evaluates reasoning quality through factual grounding, logical coherence, and answer consistency, and the Judging Reward, which adaptively determines whether deep reasoning or direct answering is appropriate. Experiments on the state-of-the-art SAIL-VL2 show that SAIL-RL improves reasoning and multimodal understanding benchmarks at both 4B and 8B scales, achieving competitive performance against commercial closed-source models such as GPT-4o, and substantially reduces hallucinations, establishing it as a principled framework for building more reliable and adaptive MLLMs. The code will be available at https://github.com/BytedanceDouyinContent/SAIL-RL.
format Preprint
id arxiv_https___arxiv_org_abs_2511_02280
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning
Shu, Fangxun
Ye, Yongjie
Liao, Yue
Kang, Zijian
Yin, Weijie
Wang, Jiacong
Liang, Xiao
Yan, Shuicheng
Feng, Chao
Computer Vision and Pattern Recognition
Computation and Language
We introduce SAIL-RL, a reinforcement learning (RL) post-training framework that enhances the reasoning capabilities of multimodal large language models (MLLMs) by teaching them when and how to think. Existing approaches are limited by outcome-only supervision, which rewards correct answers without ensuring sound reasoning, and by uniform thinking strategies, which often lead to overthinking on simple tasks and underthinking on complex ones. SAIL-RL addresses these challenges with a dual reward system: the Thinking Reward, which evaluates reasoning quality through factual grounding, logical coherence, and answer consistency, and the Judging Reward, which adaptively determines whether deep reasoning or direct answering is appropriate. Experiments on the state-of-the-art SAIL-VL2 show that SAIL-RL improves reasoning and multimodal understanding benchmarks at both 4B and 8B scales, achieving competitive performance against commercial closed-source models such as GPT-4o, and substantially reduces hallucinations, establishing it as a principled framework for building more reliable and adaptive MLLMs. The code will be available at https://github.com/BytedanceDouyinContent/SAIL-RL.
title SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2511.02280