StarVector: Generating Scalable Vector Graphics Code from Images and Text

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rodriguez, Juan A., Puri, Abhay, Agarwal, Shubham, Laradji, Issam H., Rodriguez, Pau, Rajeswar, Sai, Vazquez, David, Pal, Christopher, Pedersoli, Marco
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915314160107520
author Rodriguez, Juan A.
Puri, Abhay
Agarwal, Shubham
Laradji, Issam H.
Rodriguez, Pau
Rajeswar, Sai
Vazquez, David
Pal, Christopher
Pedersoli, Marco
author_facet Rodriguez, Juan A.
Puri, Abhay
Agarwal, Shubham
Laradji, Issam H.
Rodriguez, Pau
Rajeswar, Sai
Vazquez, David
Pal, Christopher
Pedersoli, Marco
contents Scalable Vector Graphics (SVGs) are vital for modern image rendering due to their scalability and versatility. Previous SVG generation methods have focused on curve-based vectorization, lacking semantic understanding, often producing artifacts, and struggling with SVG primitives beyond path curves. To address these issues, we introduce StarVector, a multimodal large language model for SVG generation. It performs image vectorization by understanding image semantics and using SVG primitives for compact, precise outputs. Unlike traditional methods, StarVector works directly in the SVG code space, leveraging visual understanding to apply accurate SVG primitives. To train StarVector, we create SVG-Stack, a diverse dataset of 2M samples that enables generalization across vectorization tasks and precise use of primitives like ellipses, polygons, and text. We address challenges in SVG evaluation, showing that pixel-based metrics like MSE fail to capture the unique qualities of vector graphics. We introduce SVG-Bench, a benchmark across 10 datasets, and 3 tasks: Image-to-SVG, Text-to-SVG generation, and diagram generation. Using this setup, StarVector achieves state-of-the-art performance, producing more compact and semantically rich SVGs.
format Preprint
id arxiv_https___arxiv_org_abs_2312_11556
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle StarVector: Generating Scalable Vector Graphics Code from Images and Text
Rodriguez, Juan A.
Puri, Abhay
Agarwal, Shubham
Laradji, Issam H.
Rodriguez, Pau
Rajeswar, Sai
Vazquez, David
Pal, Christopher
Pedersoli, Marco
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Scalable Vector Graphics (SVGs) are vital for modern image rendering due to their scalability and versatility. Previous SVG generation methods have focused on curve-based vectorization, lacking semantic understanding, often producing artifacts, and struggling with SVG primitives beyond path curves. To address these issues, we introduce StarVector, a multimodal large language model for SVG generation. It performs image vectorization by understanding image semantics and using SVG primitives for compact, precise outputs. Unlike traditional methods, StarVector works directly in the SVG code space, leveraging visual understanding to apply accurate SVG primitives. To train StarVector, we create SVG-Stack, a diverse dataset of 2M samples that enables generalization across vectorization tasks and precise use of primitives like ellipses, polygons, and text. We address challenges in SVG evaluation, showing that pixel-based metrics like MSE fail to capture the unique qualities of vector graphics. We introduce SVG-Bench, a benchmark across 10 datasets, and 3 tasks: Image-to-SVG, Text-to-SVG generation, and diagram generation. Using this setup, StarVector achieves state-of-the-art performance, producing more compact and semantically rich SVGs.
title StarVector: Generating Scalable Vector Graphics Code from Images and Text
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2312.11556