Rendering-Aware Reinforcement Learning for Vector Graphics Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Rodriguez, Juan A., Zhang, Haotian, Puri, Abhay, Feizi, Aarash, Pramanik, Rishav, Wichmann, Pascal, Mondal, Arnab, Samsami, Mohammad Reza, Awal, Rabiul, Taslakian, Perouz, Gella, Spandana, Rajeswar, Sai, Vazquez, David, Pal, Christopher, Pedersoli, Marco |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing
by: Rodriguez, Juan, et al.
Published: (2026)
by: Rodriguez, Juan, et al.
Published: (2026)
WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
by: Awal, Rabiul, et al.
Published: (2025)
by: Awal, Rabiul, et al.
Published: (2025)
PairBench: Are Vision-Language Models Reliable at Comparing What They See?
by: Feizi, Aarash, et al.
Published: (2025)
by: Feizi, Aarash, et al.
Published: (2025)
StarFlow: Generating Structured Workflow Outputs From Sketch Images
by: Bechard, Patrice, et al.
Published: (2025)
by: Bechard, Patrice, et al.
Published: (2025)
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
by: Jian, Xiangru, et al.
Published: (2026)
by: Jian, Xiangru, et al.
Published: (2026)
StarVector: Generating Scalable Vector Graphics Code from Images and Text
by: Rodriguez, Juan A., et al.
Published: (2023)
by: Rodriguez, Juan A., et al.
Published: (2023)
Grounding Computer Use Agents on Human Demonstrations
by: Feizi, Aarash, et al.
Published: (2025)
by: Feizi, Aarash, et al.
Published: (2025)
Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
by: Wang, Suyuchen, et al.
Published: (2025)
by: Wang, Suyuchen, et al.
Published: (2025)
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
by: Nayak, Shravan, et al.
Published: (2025)
by: Nayak, Shravan, et al.
Published: (2025)
MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
by: Gurung, Alexander, et al.
Published: (2026)
by: Gurung, Alexander, et al.
Published: (2026)
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
Mem-$π$: Adaptive Memory through Learning When and What to Generate
by: Wang, Xiaoqiang, et al.
Published: (2026)
by: Wang, Xiaoqiang, et al.
Published: (2026)
BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content
by: Monteiro, Joao, et al.
Published: (2024)
by: Monteiro, Joao, et al.
Published: (2024)
Masked Multi-Query Slot Attention for Unsupervised Object Discovery
by: Pramanik, Rishav, et al.
Published: (2024)
by: Pramanik, Rishav, et al.
Published: (2024)
InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation
by: Sahu, Gaurav, et al.
Published: (2024)
by: Sahu, Gaurav, et al.
Published: (2024)
VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
by: Zhang, Tianyu, et al.
Published: (2024)
by: Zhang, Tianyu, et al.
Published: (2024)
BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks
by: Rodriguez, Juan, et al.
Published: (2024)
by: Rodriguez, Juan, et al.
Published: (2024)
Augmenting LLM Reasoning with Dynamic Notes Writing for Complex QA
by: Maheshwary, Rishabh, et al.
Published: (2025)
by: Maheshwary, Rishabh, et al.
Published: (2025)
Efficient Dynamics Modeling in Interactive Environments with Koopman Theory
by: Mondal, Arnab Kumar, et al.
Published: (2023)
by: Mondal, Arnab Kumar, et al.
Published: (2023)
FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answering
by: Abaskohi, Amirhossein, et al.
Published: (2024)
by: Abaskohi, Amirhossein, et al.
Published: (2024)
Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding
by: Zhang, Le, et al.
Published: (2023)
by: Zhang, Le, et al.
Published: (2023)
Investigating Prompting Techniques for Zero- and Few-Shot Visual Question Answering
by: Awal, Rabiul, et al.
Published: (2023)
by: Awal, Rabiul, et al.
Published: (2023)
The Promise of RL for Autoregressive Image Editing
by: Ahmadi, Saba, et al.
Published: (2025)
by: Ahmadi, Saba, et al.
Published: (2025)
GPS-SSL: Guided Positive Sampling to Inject Prior Into Self-Supervised Learning
by: Feizi, Aarash, et al.
Published: (2024)
by: Feizi, Aarash, et al.
Published: (2024)
FairLoRA: Unpacking Bias Mitigation in Vision Models with Fairness-Driven Low-Rank Adaptation
by: Sukumaran, Rohan, et al.
Published: (2024)
by: Sukumaran, Rohan, et al.
Published: (2024)
BiXSE: Improving Dense Retrieval via Probabilistic Graded Relevance Distillation
by: Tsirigotis, Christos, et al.
Published: (2025)
by: Tsirigotis, Christos, et al.
Published: (2025)
Bezier Splatting for Fast and Differentiable Vector Graphics Rendering
by: Liu, Xi, et al.
Published: (2025)
by: Liu, Xi, et al.
Published: (2025)
Distilling Specialized Orders for Visual Generation
by: Pramanik, Rishav, et al.
Published: (2025)
by: Pramanik, Rishav, et al.
Published: (2025)
Shahname Firdausi and Transformational Management with Attitude Comparison Ulrich Transformation Pattern
by: Shirin Samsami
Published: (2015)
by: Shirin Samsami
Published: (2015)
VisMin: Visual Minimal-Change Understanding
by: Awal, Rabiul, et al.
Published: (2024)
by: Awal, Rabiul, et al.
Published: (2024)
XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
by: Monteiro, João, et al.
Published: (2024)
by: Monteiro, João, et al.
Published: (2024)
Render-in-the-Loop: Vector Graphics Generation via Visual Self-Feedback
by: Liang, Guotao, et al.
Published: (2026)
by: Liang, Guotao, et al.
Published: (2026)
Synthesis of low cost cathode electrocatalyst Pt‐Ni / C AB using DMSO as a solvent for low temperature proton exchange membrane fuel cell application
by: Abhay Pratap Singh, et al.
Published: (2024)
by: Abhay Pratap Singh, et al.
Published: (2024)
Hierarchical Retrieval at Scale: Bridging Transparency and Efficiency
by: Gupta, Shubham, et al.
Published: (2025)
by: Gupta, Shubham, et al.
Published: (2025)
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
by: Yang, Qian, et al.
Published: (2026)
by: Yang, Qian, et al.
Published: (2026)
Learning to Defer for Causal Discovery with Imperfect Experts
by: Clivio, Oscar, et al.
Published: (2025)
by: Clivio, Oscar, et al.
Published: (2025)
A Sparsity Principle for Partially Observable Causal Representation Learning
by: Xu, Danru, et al.
Published: (2024)
by: Xu, Danru, et al.
Published: (2024)
Spectral Unforgetting: Post-Hoc Recovery of Damaged Capabilities Without Retraining
by: Abro, Aarash, et al.
Published: (2026)
by: Abro, Aarash, et al.
Published: (2026)
Similar Items
-
VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing
by: Rodriguez, Juan, et al.
Published: (2026) -
WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
by: Awal, Rabiul, et al.
Published: (2025) -
PairBench: Are Vision-Language Models Reliable at Comparing What They See?
by: Feizi, Aarash, et al.
Published: (2025) -
StarFlow: Generating Structured Workflow Outputs From Sketch Images
by: Bechard, Patrice, et al.
Published: (2025) -
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
by: Jian, Xiangru, et al.
Published: (2026)