SVGFusion: A VAE-Diffusion Transformer for Vector Graphic Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866918435975331840 |
|---|---|
| author | Xing, Ximing Hu, Juncheng Xue, Ziteng Zhang, Jing Li, Buyu Wang, Sheng Xu, Dong Yu, Qian |
| author_facet | Xing, Ximing Hu, Juncheng Xue, Ziteng Zhang, Jing Li, Buyu Wang, Sheng Xu, Dong Yu, Qian |
| contents | Generating high-quality Scalable Vector Graphics (SVGs) from text remains a significant challenge. Existing LLM-based models that generate SVG code as a flat token sequence struggle with poor structural understanding and error accumulation, while optimization-based methods are slow and yield uneditable outputs. To address these limitations, we introduce SVGFusion, a unified framework that adapts the VAE-diffusion architecture to bridge the dual code-visual nature of SVGs. Our model features two core components: a Vector-Pixel Fusion Variational Autoencoder (VP-VAE) that learns a perceptually rich latent space by jointly encoding SVG code and its rendered image, and a Vector Space Diffusion Transformer (VS-DiT) that achieves globally coherent compositions through iterative refinement. Furthermore, this architecture is enhanced by a Rendering Sequence Modeling strategy, which ensures accurate object layering and occlusion. Evaluated on our novel SVGX-Dataset comprising 240k human-designed SVGs, SVGFusion establishes a new state-of-the-art, generating high-quality, editable SVGs that are strictly semantically aligned with the input text. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_10437 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | SVGFusion: A VAE-Diffusion Transformer for Vector Graphic Generation Xing, Ximing Hu, Juncheng Xue, Ziteng Zhang, Jing Li, Buyu Wang, Sheng Xu, Dong Yu, Qian Computer Vision and Pattern Recognition Graphics Machine Learning Generating high-quality Scalable Vector Graphics (SVGs) from text remains a significant challenge. Existing LLM-based models that generate SVG code as a flat token sequence struggle with poor structural understanding and error accumulation, while optimization-based methods are slow and yield uneditable outputs. To address these limitations, we introduce SVGFusion, a unified framework that adapts the VAE-diffusion architecture to bridge the dual code-visual nature of SVGs. Our model features two core components: a Vector-Pixel Fusion Variational Autoencoder (VP-VAE) that learns a perceptually rich latent space by jointly encoding SVG code and its rendered image, and a Vector Space Diffusion Transformer (VS-DiT) that achieves globally coherent compositions through iterative refinement. Furthermore, this architecture is enhanced by a Rendering Sequence Modeling strategy, which ensures accurate object layering and occlusion. Evaluated on our novel SVGX-Dataset comprising 240k human-designed SVGs, SVGFusion establishes a new state-of-the-art, generating high-quality, editable SVGs that are strictly semantically aligned with the input text. |
| title | SVGFusion: A VAE-Diffusion Transformer for Vector Graphic Generation |
| topic | Computer Vision and Pattern Recognition Graphics Machine Learning |
| url | https://arxiv.org/abs/2412.10437 |