Rendering-Aware Reinforcement Learning for Vector Graphics Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rodriguez, Juan A., Zhang, Haotian, Puri, Abhay, Feizi, Aarash, Pramanik, Rishav, Wichmann, Pascal, Mondal, Arnab, Samsami, Mohammad Reza, Awal, Rabiul, Taslakian, Perouz, Gella, Spandana, Rajeswar, Sai, Vazquez, David, Pal, Christopher, Pedersoli, Marco
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914174249992192
author Rodriguez, Juan A.
Zhang, Haotian
Puri, Abhay
Feizi, Aarash
Pramanik, Rishav
Wichmann, Pascal
Mondal, Arnab
Samsami, Mohammad Reza
Awal, Rabiul
Taslakian, Perouz
Gella, Spandana
Rajeswar, Sai
Vazquez, David
Pal, Christopher
Pedersoli, Marco
author_facet Rodriguez, Juan A.
Zhang, Haotian
Puri, Abhay
Feizi, Aarash
Pramanik, Rishav
Wichmann, Pascal
Mondal, Arnab
Samsami, Mohammad Reza
Awal, Rabiul
Taslakian, Perouz
Gella, Spandana
Rajeswar, Sai
Vazquez, David
Pal, Christopher
Pedersoli, Marco
contents Scalable Vector Graphics (SVG) offer a powerful format for representing visual designs as interpretable code. Recent advances in vision-language models (VLMs) have enabled high-quality SVG generation by framing the problem as a code generation task and leveraging large-scale pretraining. VLMs are particularly suitable for this task as they capture both global semantics and fine-grained visual patterns, while transferring knowledge across vision, natural language, and code domains. However, existing VLM approaches often struggle to produce faithful and efficient SVGs because they never observe the rendered images during training. Although differentiable rendering for autoregressive SVG code generation remains unavailable, rendered outputs can still be compared to original inputs, enabling evaluative feedback suitable for reinforcement learning (RL). We introduce RLRF (Reinforcement Learning from Rendering Feedback), an RL method that enhances SVG generation in autoregressive VLMs by leveraging feedback from rendered SVG outputs. Given an input image, the model generates SVG roll-outs that are rendered and compared to the original image to compute a reward. This visual fidelity feedback guides the model toward producing more accurate, efficient, and semantically coherent SVGs. RLRF significantly outperforms supervised fine-tuning, addressing common failure modes and enabling precise, high-quality SVG generation with strong structural understanding and generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20793
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rendering-Aware Reinforcement Learning for Vector Graphics Generation
Rodriguez, Juan A.
Zhang, Haotian
Puri, Abhay
Feizi, Aarash
Pramanik, Rishav
Wichmann, Pascal
Mondal, Arnab
Samsami, Mohammad Reza
Awal, Rabiul
Taslakian, Perouz
Gella, Spandana
Rajeswar, Sai
Vazquez, David
Pal, Christopher
Pedersoli, Marco
Computer Vision and Pattern Recognition
Artificial Intelligence
Scalable Vector Graphics (SVG) offer a powerful format for representing visual designs as interpretable code. Recent advances in vision-language models (VLMs) have enabled high-quality SVG generation by framing the problem as a code generation task and leveraging large-scale pretraining. VLMs are particularly suitable for this task as they capture both global semantics and fine-grained visual patterns, while transferring knowledge across vision, natural language, and code domains. However, existing VLM approaches often struggle to produce faithful and efficient SVGs because they never observe the rendered images during training. Although differentiable rendering for autoregressive SVG code generation remains unavailable, rendered outputs can still be compared to original inputs, enabling evaluative feedback suitable for reinforcement learning (RL). We introduce RLRF (Reinforcement Learning from Rendering Feedback), an RL method that enhances SVG generation in autoregressive VLMs by leveraging feedback from rendered SVG outputs. Given an input image, the model generates SVG roll-outs that are rendered and compared to the original image to compute a reward. This visual fidelity feedback guides the model toward producing more accurate, efficient, and semantically coherent SVGs. RLRF significantly outperforms supervised fine-tuning, addressing common failure modes and enabling precise, high-quality SVG generation with strong structural understanding and generalization.
title Rendering-Aware Reinforcement Learning for Vector Graphics Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.20793