VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912733909221376 |
|---|---|
| author | Zhao, Xuanle Jiang, Deyang Zeng, Zhixiong Chen, Lei Qiu, Haibo Huang, Jing Zhong, Yufeng Zheng, Liming Cao, Yilin Ma, Lin |
| author_facet | Zhao, Xuanle Jiang, Deyang Zeng, Zhixiong Chen, Lei Qiu, Haibo Huang, Jing Zhong, Yufeng Zheng, Liming Cao, Yilin Ma, Lin |
| contents | Multimodal code generation has garnered significant interest within the research community. Despite the notable success of recent vision-language models (VLMs) on specialized tasks like chart-to-code generation, their reliance on single-task training regimens fosters a narrow paradigm that hinders the development of generalized \textbf{VI}sio\textbf{N} \textbf{C}ode \textbf{I}ntelligence. In this work, we introduce \textbf{VinciCoder}, a unified multimodal code generation model that addresses this limitation via a two-stage training framework. We begin by constructing a large-scale Supervised Finetuning (SFT) corpus comprising 1.6M image-code pairs for tasks involving direct code generation and visual-based code refinement. Subsequently, we introduce a Visual Reinforcement Learning (ViRL) strategy, which employs a coarse-to-fine reward mechanism to improve visual fidelity by calculating visual similarity across local and global image patches. Extensive experiments on diverse multimodal code generation benchmarks demonstrate that VinciCoder achieves state-of-the-art performance, surpassing recent open-source models. The ablation study further validates the effectiveness of our proposed coarse-to-fine ViRL strategy. The data, code and model is available at https://github.com/DocTron-hub/VinciCoder. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_00391 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning Zhao, Xuanle Jiang, Deyang Zeng, Zhixiong Chen, Lei Qiu, Haibo Huang, Jing Zhong, Yufeng Zheng, Liming Cao, Yilin Ma, Lin Computer Vision and Pattern Recognition Multimodal code generation has garnered significant interest within the research community. Despite the notable success of recent vision-language models (VLMs) on specialized tasks like chart-to-code generation, their reliance on single-task training regimens fosters a narrow paradigm that hinders the development of generalized \textbf{VI}sio\textbf{N} \textbf{C}ode \textbf{I}ntelligence. In this work, we introduce \textbf{VinciCoder}, a unified multimodal code generation model that addresses this limitation via a two-stage training framework. We begin by constructing a large-scale Supervised Finetuning (SFT) corpus comprising 1.6M image-code pairs for tasks involving direct code generation and visual-based code refinement. Subsequently, we introduce a Visual Reinforcement Learning (ViRL) strategy, which employs a coarse-to-fine reward mechanism to improve visual fidelity by calculating visual similarity across local and global image patches. Extensive experiments on diverse multimodal code generation benchmarks demonstrate that VinciCoder achieves state-of-the-art performance, surpassing recent open-source models. The ablation study further validates the effectiveness of our proposed coarse-to-fine ViRL strategy. The data, code and model is available at https://github.com/DocTron-hub/VinciCoder. |
| title | VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2511.00391 |