SVGFusion: A VAE-Diffusion Transformer for Vector Graphic Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xing, Ximing, Hu, Juncheng, Xue, Ziteng, Zhang, Jing, Li, Buyu, Wang, Sheng, Xu, Dong, Yu, Qian
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918435975331840
author Xing, Ximing
Hu, Juncheng
Xue, Ziteng
Zhang, Jing
Li, Buyu
Wang, Sheng
Xu, Dong
Yu, Qian
author_facet Xing, Ximing
Hu, Juncheng
Xue, Ziteng
Zhang, Jing
Li, Buyu
Wang, Sheng
Xu, Dong
Yu, Qian
contents Generating high-quality Scalable Vector Graphics (SVGs) from text remains a significant challenge. Existing LLM-based models that generate SVG code as a flat token sequence struggle with poor structural understanding and error accumulation, while optimization-based methods are slow and yield uneditable outputs. To address these limitations, we introduce SVGFusion, a unified framework that adapts the VAE-diffusion architecture to bridge the dual code-visual nature of SVGs. Our model features two core components: a Vector-Pixel Fusion Variational Autoencoder (VP-VAE) that learns a perceptually rich latent space by jointly encoding SVG code and its rendered image, and a Vector Space Diffusion Transformer (VS-DiT) that achieves globally coherent compositions through iterative refinement. Furthermore, this architecture is enhanced by a Rendering Sequence Modeling strategy, which ensures accurate object layering and occlusion. Evaluated on our novel SVGX-Dataset comprising 240k human-designed SVGs, SVGFusion establishes a new state-of-the-art, generating high-quality, editable SVGs that are strictly semantically aligned with the input text.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10437
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SVGFusion: A VAE-Diffusion Transformer for Vector Graphic Generation
Xing, Ximing
Hu, Juncheng
Xue, Ziteng
Zhang, Jing
Li, Buyu
Wang, Sheng
Xu, Dong
Yu, Qian
Computer Vision and Pattern Recognition
Graphics
Machine Learning
Generating high-quality Scalable Vector Graphics (SVGs) from text remains a significant challenge. Existing LLM-based models that generate SVG code as a flat token sequence struggle with poor structural understanding and error accumulation, while optimization-based methods are slow and yield uneditable outputs. To address these limitations, we introduce SVGFusion, a unified framework that adapts the VAE-diffusion architecture to bridge the dual code-visual nature of SVGs. Our model features two core components: a Vector-Pixel Fusion Variational Autoencoder (VP-VAE) that learns a perceptually rich latent space by jointly encoding SVG code and its rendered image, and a Vector Space Diffusion Transformer (VS-DiT) that achieves globally coherent compositions through iterative refinement. Furthermore, this architecture is enhanced by a Rendering Sequence Modeling strategy, which ensures accurate object layering and occlusion. Evaluated on our novel SVGX-Dataset comprising 240k human-designed SVGs, SVGFusion establishes a new state-of-the-art, generating high-quality, editable SVGs that are strictly semantically aligned with the input text.
title SVGFusion: A VAE-Diffusion Transformer for Vector Graphic Generation
topic Computer Vision and Pattern Recognition
Graphics
Machine Learning
url https://arxiv.org/abs/2412.10437