VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Rodriguez, Juan, Zhang, Haotian, Puri, Abhay, Zhang, Tianyang, Pramanik, Rishav, Lin, Meng, Xie, Xiaoqing, Terral, Marco, Kaushik, Darsh, Shariff, Aly, Taslakian, Perouz, Gella, Spandana, Rajeswar, Sai, Vazquez, David, Pal, Christopher, Pedersoli, Marco |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rendering-Aware Reinforcement Learning for Vector Graphics Generation
by: Rodriguez, Juan A., et al.
Published: (2025)
by: Rodriguez, Juan A., et al.
Published: (2025)
WildSVG: Towards Reliable SVG Generation Under Real-Word Conditions
by: Terral, Marco, et al.
Published: (2026)
by: Terral, Marco, et al.
Published: (2026)
StarFlow: Generating Structured Workflow Outputs From Sketch Images
by: Bechard, Patrice, et al.
Published: (2025)
by: Bechard, Patrice, et al.
Published: (2025)
Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
by: Wang, Suyuchen, et al.
Published: (2025)
by: Wang, Suyuchen, et al.
Published: (2025)
MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
by: Gurung, Alexander, et al.
Published: (2026)
by: Gurung, Alexander, et al.
Published: (2026)
Masked Multi-Query Slot Attention for Unsupervised Object Discovery
by: Pramanik, Rishav, et al.
Published: (2024)
by: Pramanik, Rishav, et al.
Published: (2024)
Mem-$π$: Adaptive Memory through Learning When and What to Generate
by: Wang, Xiaoqiang, et al.
Published: (2026)
by: Wang, Xiaoqiang, et al.
Published: (2026)
BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
by: Awal, Rabiul, et al.
Published: (2025)
by: Awal, Rabiul, et al.
Published: (2025)
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
by: Nayak, Shravan, et al.
Published: (2025)
by: Nayak, Shravan, et al.
Published: (2025)
StarVector: Generating Scalable Vector Graphics Code from Images and Text
by: Rodriguez, Juan A., et al.
Published: (2023)
by: Rodriguez, Juan A., et al.
Published: (2023)
PairBench: Are Vision-Language Models Reliable at Comparing What They See?
by: Feizi, Aarash, et al.
Published: (2025)
by: Feizi, Aarash, et al.
Published: (2025)
VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
by: Zhang, Tianyu, et al.
Published: (2024)
by: Zhang, Tianyu, et al.
Published: (2024)
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
by: Jian, Xiangru, et al.
Published: (2026)
by: Jian, Xiangru, et al.
Published: (2026)
InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation
by: Sahu, Gaurav, et al.
Published: (2024)
by: Sahu, Gaurav, et al.
Published: (2024)
FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answering
by: Abaskohi, Amirhossein, et al.
Published: (2024)
by: Abaskohi, Amirhossein, et al.
Published: (2024)
RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content
by: Monteiro, Joao, et al.
Published: (2024)
by: Monteiro, Joao, et al.
Published: (2024)
Grounding Computer Use Agents on Human Demonstrations
by: Feizi, Aarash, et al.
Published: (2025)
by: Feizi, Aarash, et al.
Published: (2025)
Augmenting LLM Reasoning with Dynamic Notes Writing for Complex QA
by: Maheshwary, Rishabh, et al.
Published: (2025)
by: Maheshwary, Rishabh, et al.
Published: (2025)
BiXSE: Improving Dense Retrieval via Probabilistic Graded Relevance Distillation
by: Tsirigotis, Christos, et al.
Published: (2025)
by: Tsirigotis, Christos, et al.
Published: (2025)
Distilling Specialized Orders for Visual Generation
by: Pramanik, Rishav, et al.
Published: (2025)
by: Pramanik, Rishav, et al.
Published: (2025)
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
by: Yang, Qian, et al.
Published: (2026)
by: Yang, Qian, et al.
Published: (2026)
Scope: Selective Cross-modal Orchestration of Visual Perception Experts
by: Zhang, Tianyu, et al.
Published: (2025)
by: Zhang, Tianyu, et al.
Published: (2025)
DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance
by: Zhang, Peiying, et al.
Published: (2025)
by: Zhang, Peiying, et al.
Published: (2025)
Synthesis of low cost cathode electrocatalyst Pt‐Ni / C AB using DMSO as a solvent for low temperature proton exchange membrane fuel cell application
by: Abhay Pratap Singh, et al.
Published: (2024)
by: Abhay Pratap Singh, et al.
Published: (2024)
Hierarchical Retrieval at Scale: Bridging Transparency and Efficiency
by: Gupta, Shubham, et al.
Published: (2025)
by: Gupta, Shubham, et al.
Published: (2025)
Learning to Defer for Causal Discovery with Imperfect Experts
by: Clivio, Oscar, et al.
Published: (2025)
by: Clivio, Oscar, et al.
Published: (2025)
BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks
by: Rodriguez, Juan, et al.
Published: (2024)
by: Rodriguez, Juan, et al.
Published: (2024)
InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
by: Wang, Haomin, et al.
Published: (2025)
by: Wang, Haomin, et al.
Published: (2025)
ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
Unsupervised Object Discovery: A Comprehensive Survey and Unified Taxonomy
by: Villa-Vásquez, José-Fabian, et al.
Published: (2024)
by: Villa-Vásquez, José-Fabian, et al.
Published: (2024)
RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance
by: Wang, Jiuniu, et al.
Published: (2025)
by: Wang, Jiuniu, et al.
Published: (2025)
A Sparsity Principle for Partially Observable Causal Representation Learning
by: Xu, Danru, et al.
Published: (2024)
by: Xu, Danru, et al.
Published: (2024)
Multitask Kernel-based Learning with Logic Constraints
by: Diligenti, Michelangelo, et al.
Published: (2024)
by: Diligenti, Michelangelo, et al.
Published: (2024)
Data-Efficient Multitask DAgger
by: Fu, Haotian, et al.
Published: (2025)
by: Fu, Haotian, et al.
Published: (2025)
Neural Architecture Search by Learning a Hierarchical Search Space
by: Roshtkhari, Mehraveh Javan, et al.
Published: (2025)
by: Roshtkhari, Mehraveh Javan, et al.
Published: (2025)
LiveSVG: Zero-Shot SVG Animation via Video Generation
by: Levy, Matan, et al.
Published: (2026)
by: Levy, Matan, et al.
Published: (2026)
Multitask Kernel-based Learning with First-Order Logic Constraints
by: Diligenti, Michelangelo, et al.
Published: (2023)
by: Diligenti, Michelangelo, et al.
Published: (2023)
Do We Need Transformers to Play FPS Video Games?
by: Batth, Karmanbir, et al.
Published: (2025)
by: Batth, Karmanbir, et al.
Published: (2025)
Similar Items
-
Rendering-Aware Reinforcement Learning for Vector Graphics Generation
by: Rodriguez, Juan A., et al.
Published: (2025) -
WildSVG: Towards Reliable SVG Generation Under Real-Word Conditions
by: Terral, Marco, et al.
Published: (2026) -
StarFlow: Generating Structured Workflow Outputs From Sketch Images
by: Bechard, Patrice, et al.
Published: (2025) -
Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
by: Wang, Suyuchen, et al.
Published: (2025) -
MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
by: Gurung, Alexander, et al.
Published: (2026)