InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Haomin, Yin, Jinhui, Wei, Qi, Zeng, Wenguang, Gu, Lixin, Ye, Shenglong, Gao, Zhangwei, Wang, Yaohui, Zhang, Yanting, Li, Yuanqi, Guo, Yanwen, Wang, Wenhai, Chen, Kai, Qiao, Yu, Zhang, Hongjie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912889418285056
author Wang, Haomin
Yin, Jinhui
Wei, Qi
Zeng, Wenguang
Gu, Lixin
Ye, Shenglong
Gao, Zhangwei
Wang, Yaohui
Zhang, Yanting
Li, Yuanqi
Guo, Yanwen
Wang, Wenhai
Chen, Kai
Qiao, Yu
Zhang, Hongjie
author_facet Wang, Haomin
Yin, Jinhui
Wei, Qi
Zeng, Wenguang
Gu, Lixin
Ye, Shenglong
Gao, Zhangwei
Wang, Yaohui
Zhang, Yanting
Li, Yuanqi
Guo, Yanwen
Wang, Wenhai
Chen, Kai
Qiao, Yu
Zhang, Hongjie
contents General SVG modeling remains challenging due to fragmented datasets, limited transferability of methods across tasks, and the difficulty of handling structural complexity. In response, we leverage the strong transfer and generalization capabilities of multimodal large language models (MLLMs) to achieve unified modeling for SVG understanding, editing, and generation. We present the InternSVG family, an integrated data-benchmark-model suite. At its core is SAgoge, the largest and most comprehensive multimodal dataset for SVG tasks, encompassing both static graphics and dynamic animations. It covers icons, long-sequence illustrations, scientific diagrams, and dynamic animations, supporting tasks of varied difficulty levels and providing deeper hierarchies with richer attributes compared to previous datasets. Based on this resource, we introduce SArena, a companion benchmark with comprehensive task definitions and standardized evaluation that aligns with the domains and difficulty spectrum covered by SAgoge. Building on these foundations, we propose InternSVG, a unified MLLM for SVG understanding, editing, and generation with SVG-specific special tokens, subword-based embedding initialization, and a two-stage training strategy that progresses from short static SVGs to long-sequence illustrations and complex animations. This unified formulation induces positive transfer and improves overall performance. Experiments on SArena and prior benchmark confirm that InternSVG achieves substantial gains and consistently outperforms leading open and proprietary counterparts.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11341
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
Wang, Haomin
Yin, Jinhui
Wei, Qi
Zeng, Wenguang
Gu, Lixin
Ye, Shenglong
Gao, Zhangwei
Wang, Yaohui
Zhang, Yanting
Li, Yuanqi
Guo, Yanwen
Wang, Wenhai
Chen, Kai
Qiao, Yu
Zhang, Hongjie
Computer Vision and Pattern Recognition
General SVG modeling remains challenging due to fragmented datasets, limited transferability of methods across tasks, and the difficulty of handling structural complexity. In response, we leverage the strong transfer and generalization capabilities of multimodal large language models (MLLMs) to achieve unified modeling for SVG understanding, editing, and generation. We present the InternSVG family, an integrated data-benchmark-model suite. At its core is SAgoge, the largest and most comprehensive multimodal dataset for SVG tasks, encompassing both static graphics and dynamic animations. It covers icons, long-sequence illustrations, scientific diagrams, and dynamic animations, supporting tasks of varied difficulty levels and providing deeper hierarchies with richer attributes compared to previous datasets. Based on this resource, we introduce SArena, a companion benchmark with comprehensive task definitions and standardized evaluation that aligns with the domains and difficulty spectrum covered by SAgoge. Building on these foundations, we propose InternSVG, a unified MLLM for SVG understanding, editing, and generation with SVG-specific special tokens, subword-based embedding initialization, and a two-stage training strategy that progresses from short static SVGs to long-sequence illustrations and complex animations. This unified formulation induces positive transfer and improves overall performance. Experiments on SArena and prior benchmark confirm that InternSVG achieves substantial gains and consistently outperforms leading open and proprietary counterparts.
title InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.11341