LLM Code Customization with Visual Results: A Benchmark on TikZ

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Reux, Charly, Acher, Mathieu, Khelladi, Djamel Eddine, Barais, Olivier, Quinton, Clément
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908392221573120
author Reux, Charly
Acher, Mathieu
Khelladi, Djamel Eddine
Barais, Olivier
Quinton, Clément
author_facet Reux, Charly
Acher, Mathieu
Khelladi, Djamel Eddine
Barais, Olivier
Quinton, Clément
contents With the rise of AI-based code generation, customizing existing code out of natural language instructions to modify visual results -such as figures or images -has become possible, promising to reduce the need for deep programming expertise. However, even experienced developers can struggle with this task, as it requires identifying relevant code regions (feature location), generating valid code variants, and ensuring the modifications reliably align with user intent. In this paper, we introduce vTikZ, the first benchmark designed to evaluate the ability of Large Language Models (LLMs) to customize code while preserving coherent visual outcomes. Our benchmark consists of carefully curated vTikZ editing scenarios, parameterized ground truths, and a reviewing tool that leverages visual feedback to assess correctness. Empirical evaluation with stateof-the-art LLMs shows that existing solutions struggle to reliably modify code in alignment with visual intent, highlighting a gap in current AI-assisted code editing approaches. We argue that vTikZ opens new research directions for integrating LLMs with visual feedback mechanisms to improve code customization tasks in various domains beyond TikZ, including image processing, art creation, Web design, and 3D modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2505_04670
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLM Code Customization with Visual Results: A Benchmark on TikZ
Reux, Charly
Acher, Mathieu
Khelladi, Djamel Eddine
Barais, Olivier
Quinton, Clément
Software Engineering
Artificial Intelligence
With the rise of AI-based code generation, customizing existing code out of natural language instructions to modify visual results -such as figures or images -has become possible, promising to reduce the need for deep programming expertise. However, even experienced developers can struggle with this task, as it requires identifying relevant code regions (feature location), generating valid code variants, and ensuring the modifications reliably align with user intent. In this paper, we introduce vTikZ, the first benchmark designed to evaluate the ability of Large Language Models (LLMs) to customize code while preserving coherent visual outcomes. Our benchmark consists of carefully curated vTikZ editing scenarios, parameterized ground truths, and a reviewing tool that leverages visual feedback to assess correctness. Empirical evaluation with stateof-the-art LLMs shows that existing solutions struggle to reliably modify code in alignment with visual intent, highlighting a gap in current AI-assisted code editing approaches. We argue that vTikZ opens new research directions for integrating LLMs with visual feedback mechanisms to improve code customization tasks in various domains beyond TikZ, including image processing, art creation, Web design, and 3D modeling.
title LLM Code Customization with Visual Results: A Benchmark on TikZ
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2505.04670