VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Lingjie, Huang, Shaohan, Wu, Xun, Li, Yixia, Zhang, Dongdong, Wei, Furu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908488133771264
author Jiang, Lingjie
Huang, Shaohan
Wu, Xun
Li, Yixia
Zhang, Dongdong
Wei, Furu
author_facet Jiang, Lingjie
Huang, Shaohan
Wu, Xun
Li, Yixia
Zhang, Dongdong
Wei, Furu
contents Multimodal large language models (MLLMs) have significantly advanced the integration of visual and textual understanding. However, their ability to generate code from multimodal inputs remains limited. In this work, we introduce VisCodex, a unified framework that seamlessly merges vision and coding language models to empower MLLMs with strong multimodal code generation abilities. Leveraging a task vector-based model merging technique, we integrate a state-of-the-art coding LLM into a strong vision-language backbone, while preserving both visual comprehension and advanced coding skills. To support training and evaluation, we introduce the Multimodal Coding Dataset (MCD), a large-scale and diverse collection of 598k samples, including high-quality HTML code, chart image-code pairs, image-augmented StackOverflow QA, and algorithmic problems. Furthermore, we propose InfiBench-V, a novel and challenging benchmark specifically designed to assess models on visually-rich, real-world programming questions that demand a nuanced understanding of both textual and visual contexts. Extensive experiments show that VisCodex achieves state-of-the-art performance among open-source MLLMs and approaches proprietary models like GPT-4o, highlighting the effectiveness of our model merging strategy and new datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2508_09945
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
Jiang, Lingjie
Huang, Shaohan
Wu, Xun
Li, Yixia
Zhang, Dongdong
Wei, Furu
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimodal large language models (MLLMs) have significantly advanced the integration of visual and textual understanding. However, their ability to generate code from multimodal inputs remains limited. In this work, we introduce VisCodex, a unified framework that seamlessly merges vision and coding language models to empower MLLMs with strong multimodal code generation abilities. Leveraging a task vector-based model merging technique, we integrate a state-of-the-art coding LLM into a strong vision-language backbone, while preserving both visual comprehension and advanced coding skills. To support training and evaluation, we introduce the Multimodal Coding Dataset (MCD), a large-scale and diverse collection of 598k samples, including high-quality HTML code, chart image-code pairs, image-augmented StackOverflow QA, and algorithmic problems. Furthermore, we propose InfiBench-V, a novel and challenging benchmark specifically designed to assess models on visually-rich, real-world programming questions that demand a nuanced understanding of both textual and visual contexts. Extensive experiments show that VisCodex achieves state-of-the-art performance among open-source MLLMs and approaches proprietary models like GPT-4o, highlighting the effectiveness of our model merging strategy and new datasets.
title VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.09945