C2-Evo: Co-Evolving Multimodal Data and Model for Self-Improving Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Xiuwei, Hu, Wentao, Li, Hanhui, Zhou, Jun, Chen, Zisheng, Cao, Meng, Zeng, Yihan, Zhang, Kui, Yuan, Yu-Jie, Han, Jianhua, Xu, Hang, Liang, Xiaodan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915416155095040
author Chen, Xiuwei
Hu, Wentao
Li, Hanhui
Zhou, Jun
Chen, Zisheng
Cao, Meng
Zeng, Yihan
Zhang, Kui
Yuan, Yu-Jie
Han, Jianhua
Xu, Hang
Liang, Xiaodan
author_facet Chen, Xiuwei
Hu, Wentao
Li, Hanhui
Zhou, Jun
Chen, Zisheng
Cao, Meng
Zeng, Yihan
Zhang, Kui
Yuan, Yu-Jie
Han, Jianhua
Xu, Hang
Liang, Xiaodan
contents Recent advances in multimodal large language models (MLLMs) have shown impressive reasoning capabilities. However, further enhancing existing MLLMs necessitates high-quality vision-language datasets with carefully curated task complexities, which are both costly and challenging to scale. Although recent self-improving models that iteratively refine themselves offer a feasible solution, they still suffer from two core challenges: (i) most existing methods augment visual or textual data separately, resulting in discrepancies in data complexity (e.g., over-simplified diagrams paired with redundant textual descriptions); and (ii) the evolution of data and models is also separated, leading to scenarios where models are exposed to tasks with mismatched difficulty levels. To address these issues, we propose C2-Evo, an automatic, closed-loop self-improving framework that jointly evolves both training data and model capabilities. Specifically, given a base dataset and a base model, C2-Evo enhances them by a cross-modal data evolution loop and a data-model evolution loop. The former loop expands the base dataset by generating complex multimodal problems that combine structured textual sub-problems with iteratively specified geometric diagrams, while the latter loop adaptively selects the generated problems based on the performance of the base model, to conduct supervised fine-tuning and reinforcement learning alternately. Consequently, our method continuously refines its model and training data, and consistently obtains considerable performance gains across multiple mathematical reasoning benchmarks. Our code, models, and datasets will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2507_16518
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle C2-Evo: Co-Evolving Multimodal Data and Model for Self-Improving Reasoning
Chen, Xiuwei
Hu, Wentao
Li, Hanhui
Zhou, Jun
Chen, Zisheng
Cao, Meng
Zeng, Yihan
Zhang, Kui
Yuan, Yu-Jie
Han, Jianhua
Xu, Hang
Liang, Xiaodan
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Recent advances in multimodal large language models (MLLMs) have shown impressive reasoning capabilities. However, further enhancing existing MLLMs necessitates high-quality vision-language datasets with carefully curated task complexities, which are both costly and challenging to scale. Although recent self-improving models that iteratively refine themselves offer a feasible solution, they still suffer from two core challenges: (i) most existing methods augment visual or textual data separately, resulting in discrepancies in data complexity (e.g., over-simplified diagrams paired with redundant textual descriptions); and (ii) the evolution of data and models is also separated, leading to scenarios where models are exposed to tasks with mismatched difficulty levels. To address these issues, we propose C2-Evo, an automatic, closed-loop self-improving framework that jointly evolves both training data and model capabilities. Specifically, given a base dataset and a base model, C2-Evo enhances them by a cross-modal data evolution loop and a data-model evolution loop. The former loop expands the base dataset by generating complex multimodal problems that combine structured textual sub-problems with iteratively specified geometric diagrams, while the latter loop adaptively selects the generated problems based on the performance of the base model, to conduct supervised fine-tuning and reinforcement learning alternately. Consequently, our method continuously refines its model and training data, and consistently obtains considerable performance gains across multiple mathematical reasoning benchmarks. Our code, models, and datasets will be released.
title C2-Evo: Co-Evolving Multimodal Data and Model for Self-Improving Reasoning
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
url https://arxiv.org/abs/2507.16518