VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Rodriguez, Juan, Zhang, Haotian, Puri, Abhay, Zhang, Tianyang, Pramanik, Rishav, Lin, Meng, Xie, Xiaoqing, Terral, Marco, Kaushik, Darsh, Shariff, Aly, Taslakian, Perouz, Gella, Spandana, Rajeswar, Sai, Vazquez, David, Pal, Christopher, Pedersoli, Marco
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912992474431488
author Rodriguez, Juan
Zhang, Haotian
Puri, Abhay
Zhang, Tianyang
Pramanik, Rishav
Lin, Meng
Xie, Xiaoqing
Terral, Marco
Kaushik, Darsh
Shariff, Aly
Taslakian, Perouz
Gella, Spandana
Rajeswar, Sai
Vazquez, David
Pal, Christopher
Pedersoli, Marco
author_facet Rodriguez, Juan
Zhang, Haotian
Puri, Abhay
Zhang, Tianyang
Pramanik, Rishav
Lin, Meng
Xie, Xiaoqing
Terral, Marco
Kaushik, Darsh
Shariff, Aly
Taslakian, Perouz
Gella, Spandana
Rajeswar, Sai
Vazquez, David
Pal, Christopher
Pedersoli, Marco
contents We introduce VectorGym, a comprehensive benchmark suite for Scalable Vector Graphics (SVG) that spans generation from text and sketches, complex editing, and visual understanding. VectorGym addresses the lack of realistic, challenging benchmarks aligned with professional design workflows. Our benchmark comprises four tasks with expert human-authored annotations: the novel Sketch2SVG task (VG-Sketch); a new SVG editing dataset (VG-Edit) featuring complex, multi-step edits with higher-order primitives; Text2SVG generation (VG-Text); and SVG captioning (VG-Cap). Unlike prior benchmarks that rely on synthetic edits, VectorGym provides gold-standard human annotations that require semantic understanding and design intent. We also propose a multi-task reinforcement learning approach that jointly optimizes across all four tasks using rendering-based rewards. Our method, built on GRPO with curriculum learning, trains a Qwen3-VL 8B model that achieves state-of-the-art performance among open-source models, surpassing much larger models including Qwen3-VL 235B and matching GPT-4o. We also introduce a VLM-as-a-Judge metric for SVG generation, validated through human correlation studies. Our evaluation of frontier VLMs reveals significant performance gaps, positioning VectorGym as a rigorous framework for advancing visual code generation. VectorGym is publicly available on huggingface.co/datasets/ServiceNow/VectorGym.
format Preprint
id arxiv_https___arxiv_org_abs_2603_29852
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing
Rodriguez, Juan
Zhang, Haotian
Puri, Abhay
Zhang, Tianyang
Pramanik, Rishav
Lin, Meng
Xie, Xiaoqing
Terral, Marco
Kaushik, Darsh
Shariff, Aly
Taslakian, Perouz
Gella, Spandana
Rajeswar, Sai
Vazquez, David
Pal, Christopher
Pedersoli, Marco
Graphics
Artificial Intelligence
Computer Vision and Pattern Recognition
We introduce VectorGym, a comprehensive benchmark suite for Scalable Vector Graphics (SVG) that spans generation from text and sketches, complex editing, and visual understanding. VectorGym addresses the lack of realistic, challenging benchmarks aligned with professional design workflows. Our benchmark comprises four tasks with expert human-authored annotations: the novel Sketch2SVG task (VG-Sketch); a new SVG editing dataset (VG-Edit) featuring complex, multi-step edits with higher-order primitives; Text2SVG generation (VG-Text); and SVG captioning (VG-Cap). Unlike prior benchmarks that rely on synthetic edits, VectorGym provides gold-standard human annotations that require semantic understanding and design intent. We also propose a multi-task reinforcement learning approach that jointly optimizes across all four tasks using rendering-based rewards. Our method, built on GRPO with curriculum learning, trains a Qwen3-VL 8B model that achieves state-of-the-art performance among open-source models, surpassing much larger models including Qwen3-VL 235B and matching GPT-4o. We also introduce a VLM-as-a-Judge metric for SVG generation, validated through human correlation studies. Our evaluation of frontier VLMs reveals significant performance gaps, positioning VectorGym as a rigorous framework for advancing visual code generation. VectorGym is publicly available on huggingface.co/datasets/ServiceNow/VectorGym.
title VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing
topic Graphics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.29852