VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Xuanle, Jiang, Deyang, Zeng, Zhixiong, Chen, Lei, Qiu, Haibo, Huang, Jing, Zhong, Yufeng, Zheng, Liming, Cao, Yilin, Ma, Lin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912733909221376
author Zhao, Xuanle
Jiang, Deyang
Zeng, Zhixiong
Chen, Lei
Qiu, Haibo
Huang, Jing
Zhong, Yufeng
Zheng, Liming
Cao, Yilin
Ma, Lin
author_facet Zhao, Xuanle
Jiang, Deyang
Zeng, Zhixiong
Chen, Lei
Qiu, Haibo
Huang, Jing
Zhong, Yufeng
Zheng, Liming
Cao, Yilin
Ma, Lin
contents Multimodal code generation has garnered significant interest within the research community. Despite the notable success of recent vision-language models (VLMs) on specialized tasks like chart-to-code generation, their reliance on single-task training regimens fosters a narrow paradigm that hinders the development of generalized \textbf{VI}sio\textbf{N} \textbf{C}ode \textbf{I}ntelligence. In this work, we introduce \textbf{VinciCoder}, a unified multimodal code generation model that addresses this limitation via a two-stage training framework. We begin by constructing a large-scale Supervised Finetuning (SFT) corpus comprising 1.6M image-code pairs for tasks involving direct code generation and visual-based code refinement. Subsequently, we introduce a Visual Reinforcement Learning (ViRL) strategy, which employs a coarse-to-fine reward mechanism to improve visual fidelity by calculating visual similarity across local and global image patches. Extensive experiments on diverse multimodal code generation benchmarks demonstrate that VinciCoder achieves state-of-the-art performance, surpassing recent open-source models. The ablation study further validates the effectiveness of our proposed coarse-to-fine ViRL strategy. The data, code and model is available at https://github.com/DocTron-hub/VinciCoder.
format Preprint
id arxiv_https___arxiv_org_abs_2511_00391
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning
Zhao, Xuanle
Jiang, Deyang
Zeng, Zhixiong
Chen, Lei
Qiu, Haibo
Huang, Jing
Zhong, Yufeng
Zheng, Liming
Cao, Yilin
Ma, Lin
Computer Vision and Pattern Recognition
Multimodal code generation has garnered significant interest within the research community. Despite the notable success of recent vision-language models (VLMs) on specialized tasks like chart-to-code generation, their reliance on single-task training regimens fosters a narrow paradigm that hinders the development of generalized \textbf{VI}sio\textbf{N} \textbf{C}ode \textbf{I}ntelligence. In this work, we introduce \textbf{VinciCoder}, a unified multimodal code generation model that addresses this limitation via a two-stage training framework. We begin by constructing a large-scale Supervised Finetuning (SFT) corpus comprising 1.6M image-code pairs for tasks involving direct code generation and visual-based code refinement. Subsequently, we introduce a Visual Reinforcement Learning (ViRL) strategy, which employs a coarse-to-fine reward mechanism to improve visual fidelity by calculating visual similarity across local and global image patches. Extensive experiments on diverse multimodal code generation benchmarks demonstrate that VinciCoder achieves state-of-the-art performance, surpassing recent open-source models. The ablation study further validates the effectiveness of our proposed coarse-to-fine ViRL strategy. The data, code and model is available at https://github.com/DocTron-hub/VinciCoder.
title VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.00391