R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jingyi, Lin, Tianyi, Yao, Huanjin, Lan, Xiang, Liu, Shunyu, Huang, Jiaxing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912872029749248
author Zhang, Jingyi
Lin, Tianyi
Yao, Huanjin
Lan, Xiang
Liu, Shunyu
Huang, Jiaxing
author_facet Zhang, Jingyi
Lin, Tianyi
Yao, Huanjin
Lan, Xiang
Liu, Shunyu
Huang, Jiaxing
contents In this work, we aim to develop effective data synthesis techniques that autonomously synthesize multimodal training data for enhancing MLLMs in solving complex real-world tasks. To this end, we propose Collective Adversarial Data Synthesis (CADS), a novel and general approach to synthesize high-quality, diverse and challenging multimodal data for MLLMs. The core idea of CADS is to leverage collective intelligence to ensure high-quality and diverse generation, while exploring adversarial learning to synthesize challenging samples for effectively driving model improvement. Specifically, CADS operates with two cyclic phases, i.e., Collective Adversarial Data Generation (CAD-Generate) and Collective Adversarial Data Judgment (CAD-Judge). CAD-Generate leverages collective knowledge to jointly generate new and diverse multimodal data, while CAD-Judge collaboratively assesses the quality of synthesized data. In addition, CADS introduces an Adversarial Context Optimization mechanism to optimize the generation context to encourage challenging and high-value data generation. With CADS, we construct MMSynthetic-20K and train our model R1-SyntheticVL, which demonstrates superior performance on various benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03300
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model?
Zhang, Jingyi
Lin, Tianyi
Yao, Huanjin
Lan, Xiang
Liu, Shunyu
Huang, Jiaxing
Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
In this work, we aim to develop effective data synthesis techniques that autonomously synthesize multimodal training data for enhancing MLLMs in solving complex real-world tasks. To this end, we propose Collective Adversarial Data Synthesis (CADS), a novel and general approach to synthesize high-quality, diverse and challenging multimodal data for MLLMs. The core idea of CADS is to leverage collective intelligence to ensure high-quality and diverse generation, while exploring adversarial learning to synthesize challenging samples for effectively driving model improvement. Specifically, CADS operates with two cyclic phases, i.e., Collective Adversarial Data Generation (CAD-Generate) and Collective Adversarial Data Judgment (CAD-Judge). CAD-Generate leverages collective knowledge to jointly generate new and diverse multimodal data, while CAD-Judge collaboratively assesses the quality of synthesized data. In addition, CADS introduces an Adversarial Context Optimization mechanism to optimize the generation context to encourage challenging and high-value data generation. With CADS, we construct MMSynthetic-20K and train our model R1-SyntheticVL, which demonstrates superior performance on various benchmarks.
title R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model?
topic Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.03300