RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Jiuniu, Zhang, Gongjie, Qian, Quanhao, Gao, Junlong, Zhao, Deli, Xu, Ran
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912671760121856
author Wang, Jiuniu
Zhang, Gongjie
Qian, Quanhao
Gao, Junlong
Zhao, Deli
Xu, Ran
author_facet Wang, Jiuniu
Zhang, Gongjie
Qian, Quanhao
Gao, Junlong
Zhao, Deli
Xu, Ran
contents Scalable Vector Graphics (SVGs) are fundamental to digital design and robot control, encoding not only visual structure but also motion paths in interactive drawings. In this work, we introduce RoboSVG, a unified multimodal framework for generating interactive SVGs guided by textual, visual, and numerical signals. Given an input query, the RoboSVG model first produces multimodal guidance, then synthesizes candidate SVGs through dedicated generation modules, and finally refines them under numerical guidance to yield high-quality outputs. To support this framework, we construct RoboDraw, a large-scale dataset of one million examples, each pairing an SVG generation condition (e.g., text, image, and partial SVG) with its corresponding ground-truth SVG code. RoboDraw dataset enables systematic study of four tasks, including basic generation (Text-to-SVG, Image-to-SVG) and interactive generation (PartialSVG-to-SVG, PartialImage-to-SVG). Extensive experiments demonstrate that RoboSVG achieves superior query compliance and visual fidelity across tasks, establishing a new state of the art in versatile SVG generation. The dataset and source code of this project will be publicly available soon.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22684
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance
Wang, Jiuniu
Zhang, Gongjie
Qian, Quanhao
Gao, Junlong
Zhao, Deli
Xu, Ran
Computer Vision and Pattern Recognition
Computation and Language
Scalable Vector Graphics (SVGs) are fundamental to digital design and robot control, encoding not only visual structure but also motion paths in interactive drawings. In this work, we introduce RoboSVG, a unified multimodal framework for generating interactive SVGs guided by textual, visual, and numerical signals. Given an input query, the RoboSVG model first produces multimodal guidance, then synthesizes candidate SVGs through dedicated generation modules, and finally refines them under numerical guidance to yield high-quality outputs. To support this framework, we construct RoboDraw, a large-scale dataset of one million examples, each pairing an SVG generation condition (e.g., text, image, and partial SVG) with its corresponding ground-truth SVG code. RoboDraw dataset enables systematic study of four tasks, including basic generation (Text-to-SVG, Image-to-SVG) and interactive generation (PartialSVG-to-SVG, PartialImage-to-SVG). Extensive experiments demonstrate that RoboSVG achieves superior query compliance and visual fidelity across tasks, establishing a new state of the art in versatile SVG generation. The dataset and source code of this project will be publicly available soon.
title RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2510.22684