DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Peiying, Zhao, Nanxuan, Fisher, Matthew, Xu, Yiran, Liao, Jing, Liu, Difan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915669837086720
author Zhang, Peiying
Zhao, Nanxuan
Fisher, Matthew
Xu, Yiran
Liao, Jing
Liu, Difan
author_facet Zhang, Peiying
Zhao, Nanxuan
Fisher, Matthew
Xu, Yiran
Liao, Jing
Liu, Difan
contents Recent vision-language model (VLM)-based approaches have achieved impressive results on SVG generation. However, because they generate only text and lack visual signals during decoding, they often struggle with complex semantics and fail to produce visually appealing or geometrically coherent SVGs. We introduce DuetSVG, a unified multimodal model that jointly generates image tokens and corresponding SVG tokens in an end-to-end manner. DuetSVG is trained on both image and SVG datasets. At inference, we apply a novel test-time scaling strategy that leverages the model's native visual predictions as guidance to improve SVG decoding quality. Extensive experiments show that our method outperforms existing methods, producing visually faithful, semantically aligned, and syntactically clean SVGs across a wide range of applications.
format Preprint
id arxiv_https___arxiv_org_abs_2512_10894
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance
Zhang, Peiying
Zhao, Nanxuan
Fisher, Matthew
Xu, Yiran
Liao, Jing
Liu, Difan
Computer Vision and Pattern Recognition
Recent vision-language model (VLM)-based approaches have achieved impressive results on SVG generation. However, because they generate only text and lack visual signals during decoding, they often struggle with complex semantics and fail to produce visually appealing or geometrically coherent SVGs. We introduce DuetSVG, a unified multimodal model that jointly generates image tokens and corresponding SVG tokens in an end-to-end manner. DuetSVG is trained on both image and SVG datasets. At inference, we apply a novel test-time scaling strategy that leverages the model's native visual predictions as guidance to improve SVG decoding quality. Extensive experiments show that our method outperforms existing methods, producing visually faithful, semantically aligned, and syntactically clean SVGs across a wide range of applications.
title DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.10894