ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Chenchen, Li, Yuhang, Xu, Can, Liu, Jiaheng, Liu, Ao, Zhou, Changzhi, Deng, Ken, Wu, Dengpeng, Huang, Guanhua, Li, Kejiao, Yi, Qi, Xiong, Ruibin, Hu, Shihui, Zhang, Yue, Jiang, Yuhao, Xu, Zenan, Zhang, Yuanxing, Zhou, Wiggin, Zhou, Chayse, Lian, Fengzong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918150271926272
author Zhang, Chenchen
Li, Yuhang
Xu, Can
Liu, Jiaheng
Liu, Ao
Zhou, Changzhi
Deng, Ken
Wu, Dengpeng
Huang, Guanhua
Li, Kejiao
Yi, Qi
Xiong, Ruibin
Hu, Shihui
Zhang, Yue
Jiang, Yuhao
Xu, Zenan
Zhang, Yuanxing
Zhou, Wiggin
Zhou, Chayse
Lian, Fengzong
author_facet Zhang, Chenchen
Li, Yuhang
Xu, Can
Liu, Jiaheng
Liu, Ao
Zhou, Changzhi
Deng, Ken
Wu, Dengpeng
Huang, Guanhua
Li, Kejiao
Yi, Qi
Xiong, Ruibin
Hu, Shihui
Zhang, Yue
Jiang, Yuhao
Xu, Zenan
Zhang, Yuanxing
Zhou, Wiggin
Zhou, Chayse
Lian, Fengzong
contents The generative capabilities of Large Language Models (LLMs) are rapidly expanding from static code to dynamic, interactive visual artifacts. This progress is bottlenecked by a critical evaluation gap: established benchmarks focus on algorithmic correctness and are blind to the visual fidelity and interactive integrity that define modern user experiences. To bridge this gap, we introduce ArtifactsBench, a new benchmark and paradigm for the automated, multimodal evaluation of visual code generation. Our framework programmatically renders each generated artifact and captures its dynamic behavior through temporal screenshots. This visual evidence, alongside the source code, is then assessed by a Multimodal LLM (MLLM)-as-Judge, which is rigorously guided by a fine-grained, per-task checklist to ensure holistic and reproducible scoring. We construct a new benchmark of 1,825 diverse tasks and evaluate over 30 leading LLMs. Our automated evaluation achieves a striking 94.4% ranking consistency with WebDev Arena, the gold-standard for human preference in web development, and over 90% pairwise agreement with human experts. This establishes ArtifactsBench as the first framework to reliably automate the assessment of human-perceived quality at scale. Our analysis provides a high-resolution map of the current SOTA, revealing that generalist models often outperform domain-specific ones. We open-source ArtifactsBench, including the benchmark, evaluation harness, and baseline results at https://artifactsbenchmark.github.io/, to provide the community with a scalable and accurate tool to accelerate the development of user-centric generative models.
format Preprint
id arxiv_https___arxiv_org_abs_2507_04952
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
Zhang, Chenchen
Li, Yuhang
Xu, Can
Liu, Jiaheng
Liu, Ao
Zhou, Changzhi
Deng, Ken
Wu, Dengpeng
Huang, Guanhua
Li, Kejiao
Yi, Qi
Xiong, Ruibin
Hu, Shihui
Zhang, Yue
Jiang, Yuhao
Xu, Zenan
Zhang, Yuanxing
Zhou, Wiggin
Zhou, Chayse
Lian, Fengzong
Computation and Language
Software Engineering
The generative capabilities of Large Language Models (LLMs) are rapidly expanding from static code to dynamic, interactive visual artifacts. This progress is bottlenecked by a critical evaluation gap: established benchmarks focus on algorithmic correctness and are blind to the visual fidelity and interactive integrity that define modern user experiences. To bridge this gap, we introduce ArtifactsBench, a new benchmark and paradigm for the automated, multimodal evaluation of visual code generation. Our framework programmatically renders each generated artifact and captures its dynamic behavior through temporal screenshots. This visual evidence, alongside the source code, is then assessed by a Multimodal LLM (MLLM)-as-Judge, which is rigorously guided by a fine-grained, per-task checklist to ensure holistic and reproducible scoring. We construct a new benchmark of 1,825 diverse tasks and evaluate over 30 leading LLMs. Our automated evaluation achieves a striking 94.4% ranking consistency with WebDev Arena, the gold-standard for human preference in web development, and over 90% pairwise agreement with human experts. This establishes ArtifactsBench as the first framework to reliably automate the assessment of human-perceived quality at scale. Our analysis provides a high-resolution map of the current SOTA, revealing that generalist models often outperform domain-specific ones. We open-source ArtifactsBench, including the benchmark, evaluation harness, and baseline results at https://artifactsbenchmark.github.io/, to provide the community with a scalable and accurate tool to accelerate the development of user-centric generative models.
title ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
topic Computation and Language
Software Engineering
url https://arxiv.org/abs/2507.04952