Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Sudong, Huang, Weiquan, Yu, Xiaomin, Yang, Zuhao, Lin, Hehai, Wu, Keming, Xiao, Chaojun, Chen, Chen, Wang, Wenxuan, Zhu, Beier, Zhang, Yunjian, Qin, Chengwei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913082418135040
author Wang, Sudong
Huang, Weiquan
Yu, Xiaomin
Yang, Zuhao
Lin, Hehai
Wu, Keming
Xiao, Chaojun
Chen, Chen
Wang, Wenxuan
Zhu, Beier
Zhang, Yunjian
Qin, Chengwei
author_facet Wang, Sudong
Huang, Weiquan
Yu, Xiaomin
Yang, Zuhao
Lin, Hehai
Wu, Keming
Xiao, Chaojun
Chen, Chen
Wang, Wenxuan
Zhu, Beier
Zhang, Yunjian
Qin, Chengwei
contents The standard post-training recipe for large multimodal models (LMMs) applies supervised fine-tuning (SFT) on curated demonstrations followed by reinforcement learning with verifiable rewards (RLVR). However, SFT introduces distributional drift that neither preserves the model's original capabilities nor faithfully matches the supervision distribution. This problem is further amplified in multimodal reasoning, where perception errors and reasoning failures follow distinct drift patterns that compound during subsequent RL. We introduce PRISM, a three-stage pipeline that mitigates this drift by inserting an explicit distribution-alignment stage between SFT and RLVR. Building on the principle of on-policy distillation (OPD), PRISM casts alignment as a black-box, response-level adversarial game between the policy and a Mixture-of-Experts (MoE) discriminator with dedicated perception and reasoning experts, providing disentangled corrective signals that steer the policy toward the supervision distribution without requiring access to teacher logits. While 1.26M public demonstrations suffice for broad SFT initialization, distribution alignment demands higher-fidelity supervision; we therefore curate 113K additional demonstrations from Gemini 3 Flash, featuring dense visual grounding and step-by-step reasoning on the hardest unsolved problems. Experiments on Qwen3-VL show that PRISM consistently improves downstream RLVR performance across multiple RL algorithms (GRPO, DAPO, GSPO) and diverse multimodal benchmarks, improving average accuracy by +4.4 and +6.0 points over the SFT-to-RLVR baseline on 4B and 8B, respectively. Our code, data, and model checkpoints are publicly available at https://github.com/XIAO4579/PRISM.
format Preprint
id arxiv_https___arxiv_org_abs_2604_28123
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL
Wang, Sudong
Huang, Weiquan
Yu, Xiaomin
Yang, Zuhao
Lin, Hehai
Wu, Keming
Xiao, Chaojun
Chen, Chen
Wang, Wenxuan
Zhu, Beier
Zhang, Yunjian
Qin, Chengwei
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
The standard post-training recipe for large multimodal models (LMMs) applies supervised fine-tuning (SFT) on curated demonstrations followed by reinforcement learning with verifiable rewards (RLVR). However, SFT introduces distributional drift that neither preserves the model's original capabilities nor faithfully matches the supervision distribution. This problem is further amplified in multimodal reasoning, where perception errors and reasoning failures follow distinct drift patterns that compound during subsequent RL. We introduce PRISM, a three-stage pipeline that mitigates this drift by inserting an explicit distribution-alignment stage between SFT and RLVR. Building on the principle of on-policy distillation (OPD), PRISM casts alignment as a black-box, response-level adversarial game between the policy and a Mixture-of-Experts (MoE) discriminator with dedicated perception and reasoning experts, providing disentangled corrective signals that steer the policy toward the supervision distribution without requiring access to teacher logits. While 1.26M public demonstrations suffice for broad SFT initialization, distribution alignment demands higher-fidelity supervision; we therefore curate 113K additional demonstrations from Gemini 3 Flash, featuring dense visual grounding and step-by-step reasoning on the hardest unsolved problems. Experiments on Qwen3-VL show that PRISM consistently improves downstream RLVR performance across multiple RL algorithms (GRPO, DAPO, GSPO) and diverse multimodal benchmarks, improving average accuracy by +4.4 and +6.0 points over the SFT-to-RLVR baseline on 4B and 8B, respectively. Our code, data, and model checkpoints are publicly available at https://github.com/XIAO4579/PRISM.
title Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2604.28123