LLM Code Customization with Visual Results: A Benchmark on TikZ
Fuente:
arXiv
Saved in:
| Main Authors: | Reux, Charly, Acher, Mathieu, Khelladi, Djamel Eddine, Barais, Olivier, Quinton, Clément |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unify and Triumph: Polyglot, Diverse, and Self-Consistent Generation of Unit Tests with LLMs
by: Khelladi, Djamel Eddine, et al.
Published: (2025)
by: Khelladi, Djamel Eddine, et al.
Published: (2025)
SpaceTime Programming: Live and Omniscient Exploration of Code and Execution
by: Döderlein, Jean-Baptiste, et al.
Published: (2026)
by: Döderlein, Jean-Baptiste, et al.
Published: (2026)
Linux Kernel Configurations at Scale: A Dataset for Performance and Evolution Analysis
by: Borges, Heraldo, et al.
Published: (2025)
by: Borges, Heraldo, et al.
Published: (2025)
A Performance Study of LLM-Generated Code on Leetcode
by: Coignion, Tristan, et al.
Published: (2024)
by: Coignion, Tristan, et al.
Published: (2024)
Did You Forkget It? Detecting One-Day Vulnerabilities in Open-source ForksWith Global History Analysis
by: Lefeuvre, Romain, et al.
Published: (2025)
by: Lefeuvre, Romain, et al.
Published: (2025)
Piloting Copilot, Codex, and StarCoder2: Hot Temperature, Cold Prompts, or Black Magic?
by: Döderlein, Jean-Baptiste, et al.
Published: (2022)
by: Döderlein, Jean-Baptiste, et al.
Published: (2022)
Green My LLM: Studying the key factors affecting the energy consumption of code assistants
by: Coignion, Tristan, et al.
Published: (2024)
by: Coignion, Tristan, et al.
Published: (2024)
Exploring Performance Trade-offs in JHipster
by: Guégain, Edouard, et al.
Published: (2024)
by: Guégain, Edouard, et al.
Published: (2024)
EvoCodeBench: A Human-Performance Benchmark for Self-Evolving LLM-Driven Coding Systems
by: Zhang, Wentao, et al.
Published: (2026)
by: Zhang, Wentao, et al.
Published: (2026)
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
by: Pan, Zhiyuan, et al.
Published: (2025)
by: Pan, Zhiyuan, et al.
Published: (2025)
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
by: Cui, Yi
Published: (2025)
by: Cui, Yi
Published: (2025)
Does Code Cleanliness Affect Coding Agents? A Controlled Minimal-Pair Study
by: Trivedi, Priyansh, et al.
Published: (2026)
by: Trivedi, Priyansh, et al.
Published: (2026)
Toward Automated and Trustworthy Scientific Analysis and Visualization with LLM-Generated Code
by: Chakroborti, Apu Kumar, et al.
Published: (2025)
by: Chakroborti, Apu Kumar, et al.
Published: (2025)
CodeFuse-CommitEval: Towards Benchmarking LLM's Power on Commit Message and Code Change Inconsistency Detection
by: Zhang, Qingyu, et al.
Published: (2025)
by: Zhang, Qingyu, et al.
Published: (2025)
1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World
by: Xu, Qiao, et al.
Published: (2026)
by: Xu, Qiao, et al.
Published: (2026)
Automated Customization of LLMs for Enterprise Code Repositories Using Semantic Scopes
by: Finkler, Ulrich, et al.
Published: (2026)
by: Finkler, Ulrich, et al.
Published: (2026)
RustEvo^2: An Evolving Benchmark for API Evolution in LLM-based Rust Code Generation
by: Liang, Linxi, et al.
Published: (2025)
by: Liang, Linxi, et al.
Published: (2025)
A Benchmark for Localizing Code and Non-Code Issues in Software Projects
by: Zhang, Zejun, et al.
Published: (2025)
by: Zhang, Zejun, et al.
Published: (2025)
RedCode: Risky Code Execution and Generation Benchmark for Code Agents
by: Guo, Chengquan, et al.
Published: (2024)
by: Guo, Chengquan, et al.
Published: (2024)
Code Review Agent Benchmark
by: Zhang, Yuntong, et al.
Published: (2026)
by: Zhang, Yuntong, et al.
Published: (2026)
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
by: Chu, Junjie, et al.
Published: (2026)
by: Chu, Junjie, et al.
Published: (2026)
AXIOM: Benchmarking LLM-as-a-Judge for Code via Rule-Based Perturbation and Multisource Quality Calibration
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
by: Jiang, Hongchao, et al.
Published: (2025)
by: Jiang, Hongchao, et al.
Published: (2025)
Lyra: A Benchmark for Turducken-Style Code Generation
by: Liang, Qingyuan, et al.
Published: (2021)
by: Liang, Qingyuan, et al.
Published: (2021)
Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code Rewriting
by: Ye, Tong, et al.
Published: (2024)
by: Ye, Tong, et al.
Published: (2024)
How Many Tries Does It Take? Iterative Self-Repair in LLM Code Generation Across Model Scales and Benchmarks
by: Arimbur, Johin Johny
Published: (2026)
by: Arimbur, Johin Johny
Published: (2026)
Benchmarking LLM Code Generation for Audio Programming with Visual Dataflow Languages
by: Zhang, William, et al.
Published: (2024)
by: Zhang, William, et al.
Published: (2024)
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
by: Syromiatnikov, Mykyta, et al.
Published: (2025)
by: Syromiatnikov, Mykyta, et al.
Published: (2025)
A New Benchmark for the Appropriate Evaluation of RTL Code Optimization
by: Lu, Yao, et al.
Published: (2026)
by: Lu, Yao, et al.
Published: (2026)
Benchmarks and Metrics for Evaluations of Code Generation: A Critical Review
by: Paul, Debalina Ghosh, et al.
Published: (2024)
by: Paul, Debalina Ghosh, et al.
Published: (2024)
Beyond Retrieval: A Multitask Benchmark and Model for Code Search
by: Xue, Siqiao, et al.
Published: (2026)
by: Xue, Siqiao, et al.
Published: (2026)
SWE Context Bench: A Benchmark for Context Learning in Coding
by: Zhu, Jiayuan, et al.
Published: (2026)
by: Zhu, Jiayuan, et al.
Published: (2026)
From Charts to Code: A Hierarchical Benchmark for Multimodal Models
by: Tang, Jiahao, et al.
Published: (2025)
by: Tang, Jiahao, et al.
Published: (2025)
SWE-QA: A Dataset and Benchmark for Complex Code Understanding
by: Elkoussy, Laïla, et al.
Published: (2026)
by: Elkoussy, Laïla, et al.
Published: (2026)
CodeSense: a Real-World Benchmark and Dataset for Code Semantic Reasoning
by: Roy, Monoshi Kumar, et al.
Published: (2025)
by: Roy, Monoshi Kumar, et al.
Published: (2025)
Efficiently Ranking Software Variants with Minimal Benchmarks
by: Matricon, Théo, et al.
Published: (2025)
by: Matricon, Théo, et al.
Published: (2025)
Specification and Detection of LLM Code Smells
by: Mahmoudi, Brahim, et al.
Published: (2025)
by: Mahmoudi, Brahim, et al.
Published: (2025)
Investigating The Smells of LLM Generated Code
by: Paul, Debalina Ghosh, et al.
Published: (2025)
by: Paul, Debalina Ghosh, et al.
Published: (2025)
Prompting for Performance: Exploring LLMs for Configuring Software
by: Spieker, Helge, et al.
Published: (2025)
by: Spieker, Helge, et al.
Published: (2025)
Code Copycat Conundrum: Demystifying Repetition in LLM-based Code Generation
by: Liu, Mingwei, et al.
Published: (2025)
by: Liu, Mingwei, et al.
Published: (2025)
Similar Items
-
Unify and Triumph: Polyglot, Diverse, and Self-Consistent Generation of Unit Tests with LLMs
by: Khelladi, Djamel Eddine, et al.
Published: (2025) -
SpaceTime Programming: Live and Omniscient Exploration of Code and Execution
by: Döderlein, Jean-Baptiste, et al.
Published: (2026) -
Linux Kernel Configurations at Scale: A Dataset for Performance and Evolution Analysis
by: Borges, Heraldo, et al.
Published: (2025) -
A Performance Study of LLM-Generated Code on Leetcode
by: Coignion, Tristan, et al.
Published: (2024) -
Did You Forkget It? Detecting One-Day Vulnerabilities in Open-source ForksWith Global History Analysis
by: Lefeuvre, Romain, et al.
Published: (2025)