Unlocking the Potential of Difficulty Prior in RL-based Multimodal Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Mingrui, Liu, Haogeng, Liang, Hao, Huang, Huaibo, Zhang, Wentao, He, Ran
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912760626937856
author Chen, Mingrui
Liu, Haogeng
Liang, Hao
Huang, Huaibo
Zhang, Wentao
He, Ran
author_facet Chen, Mingrui
Liu, Haogeng
Liang, Hao
Huang, Huaibo
Zhang, Wentao
He, Ran
contents In this work, we investigate how explicitly modeling problem's difficulty prior information shapes the effectiveness of reinforcement learning based fine-tuning for multimodal reasoning. Our exploration mainly comprises of following three perspective: First, through offline data curation, we analyze the U-shaped difficulty distribution of two given datasets using the base model by multi-round sampling, and then filter out prompts that are either too simple or extremely difficult to provide meaningful gradients and perform subsequent two-stage training. Second, we implement an online advantage differentiation, computing group-wise empirical accuracy as a difficulty proxy to adaptively reweight advantages estimation, providing stronger learning signals for more challenging problems. Finally, we introduce difficulty hints as explicit prompts for more complex samples in the second training stage, encouraging the model to calibrate its reasoning depth and perform reflective validation checks. Our comprehensive approach demonstrates significant performances across various multi-modal mathematical reasoning benchmarks with only 2K+0.6K two-stage training data.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13261
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unlocking the Potential of Difficulty Prior in RL-based Multimodal Reasoning
Chen, Mingrui
Liu, Haogeng
Liang, Hao
Huang, Huaibo
Zhang, Wentao
He, Ran
Computer Vision and Pattern Recognition
In this work, we investigate how explicitly modeling problem's difficulty prior information shapes the effectiveness of reinforcement learning based fine-tuning for multimodal reasoning. Our exploration mainly comprises of following three perspective: First, through offline data curation, we analyze the U-shaped difficulty distribution of two given datasets using the base model by multi-round sampling, and then filter out prompts that are either too simple or extremely difficult to provide meaningful gradients and perform subsequent two-stage training. Second, we implement an online advantage differentiation, computing group-wise empirical accuracy as a difficulty proxy to adaptively reweight advantages estimation, providing stronger learning signals for more challenging problems. Finally, we introduce difficulty hints as explicit prompts for more complex samples in the second training stage, encouraging the model to calibrate its reasoning depth and perform reflective validation checks. Our comprehensive approach demonstrates significant performances across various multi-modal mathematical reasoning benchmarks with only 2K+0.6K two-stage training data.
title Unlocking the Potential of Difficulty Prior in RL-based Multimodal Reasoning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.13261