Saved in:
Bibliographic Details
Main Authors: Zhu, Hongda, Zhang, Yiwen, Zhao, Bing, Ding, Jingzhe, Liu, Siyao, Liu, Tong, Wang, Dandan, Liu, Yanan, Li, Zhaojian
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2506.13832
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908412481110016
author Zhu, Hongda
Zhang, Yiwen
Zhao, Bing
Ding, Jingzhe
Liu, Siyao
Liu, Tong
Wang, Dandan
Liu, Yanan
Li, Zhaojian
author_facet Zhu, Hongda
Zhang, Yiwen
Zhao, Bing
Ding, Jingzhe
Liu, Siyao
Liu, Tong
Wang, Dandan
Liu, Yanan
Li, Zhaojian
contents Large Language Models (LLMs) have made significant strides in front-end code generation. However, existing benchmarks exhibit several critical limitations: many tasks are overly simplistic, test cases often lack rigor, and end-to-end validation is absent. These issues hinder the accurate assessment of model performance. To address these challenges, we present FrontendBench, a benchmark co-developed by humans and LLMs. FrontendBench categorizes tasks based on code functionality and incorporates interactive test scenarios, enabling a more comprehensive and practical evaluation of front-end code generation capabilities. The benchmark comprises 148 meticulously crafted prompt-test case pairs spanning five levels of web components, from basic UI elements to complex interactive features. Each task reflects realistic front-end development challenges. Furthermore, we introduce an automatic evaluation framework that executes generated code within a sandbox environment and assesses outcomes using predefined test scripts. This framework achieves a 90.54% agreement rate with expert human evaluations, demonstrating high reliability. We benchmark several state-of-the-art LLMs on FrontendBench and observe substantial performance disparities in handling real-world front-end tasks. These results highlight FrontendBench as a reliable and scalable benchmark, supporting consistent multimodal evaluation and providing a robust foundation for future research in front-end code generation. Our data and code will be released soon.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13832
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
Zhu, Hongda
Zhang, Yiwen
Zhao, Bing
Ding, Jingzhe
Liu, Siyao
Liu, Tong
Wang, Dandan
Liu, Yanan
Li, Zhaojian
Software Engineering
Artificial Intelligence
Large Language Models (LLMs) have made significant strides in front-end code generation. However, existing benchmarks exhibit several critical limitations: many tasks are overly simplistic, test cases often lack rigor, and end-to-end validation is absent. These issues hinder the accurate assessment of model performance. To address these challenges, we present FrontendBench, a benchmark co-developed by humans and LLMs. FrontendBench categorizes tasks based on code functionality and incorporates interactive test scenarios, enabling a more comprehensive and practical evaluation of front-end code generation capabilities. The benchmark comprises 148 meticulously crafted prompt-test case pairs spanning five levels of web components, from basic UI elements to complex interactive features. Each task reflects realistic front-end development challenges. Furthermore, we introduce an automatic evaluation framework that executes generated code within a sandbox environment and assesses outcomes using predefined test scripts. This framework achieves a 90.54% agreement rate with expert human evaluations, demonstrating high reliability. We benchmark several state-of-the-art LLMs on FrontendBench and observe substantial performance disparities in handling real-world front-end tasks. These results highlight FrontendBench as a reliable and scalable benchmark, supporting consistent multimodal evaluation and providing a robust foundation for future research in front-end code generation. Our data and code will be released soon.
title FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2506.13832