Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Fan, Dong, Lishuai, Gao, Cuiyun, Chen, Yujia, Huang, Yiming, Xiao, Yang, Liao, Qing
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913174744203264
author Wu, Fan
Dong, Lishuai
Gao, Cuiyun
Chen, Yujia
Huang, Yiming
Xiao, Yang
Liao, Qing
author_facet Wu, Fan
Dong, Lishuai
Gao, Cuiyun
Chen, Yujia
Huang, Yiming
Xiao, Yang
Liao, Qing
contents Recent advancements in multimodal large language models (MLLMs) have achieved remarkable progress in multimodal reasoning and code generation, catalyzing a new paradigm for front-end development. In particular, these models can directly transform visual designs into executable code, significantly improving the efficiency and adaptability of web development. Modern web applications are dynamic and interactive, featuring frequent user-page interactions. However, existing benchmarks largely evaluate the code generation of static webpages, ignoring the complex interactive behaviors in real-world applications. Besides, their evaluation criteria remain confined to visual fidelity and code structure, overlooking the interaction consistency between the generated and the reference webpages. To address these limitations, we introduce WebIGBench, the first benchmark designed to evaluate code generation for interactive webpages with complex interactions. By combining manually designed interaction paths with UI automation, we collected 103 complex webpages from real-world websites. This benchmark covers 5 popular interactive action types (e.g., click, input) involving 871 distinct interactive actions. Moreover, we propose a novel evaluation pipeline to address the gap in automated assessment of interactive actions. Extensive experiments on several representative MLLMs reveal the performance boundaries of current models in interactive webpage code generation using WebIGBench. The proposed benchmark is available at https://github.com/anoa12159-hue/WebIGBench_eval.
format Preprint
id arxiv_https___arxiv_org_abs_2606_00154
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages
Wu, Fan
Dong, Lishuai
Gao, Cuiyun
Chen, Yujia
Huang, Yiming
Xiao, Yang
Liao, Qing
Software Engineering
Artificial Intelligence
Recent advancements in multimodal large language models (MLLMs) have achieved remarkable progress in multimodal reasoning and code generation, catalyzing a new paradigm for front-end development. In particular, these models can directly transform visual designs into executable code, significantly improving the efficiency and adaptability of web development. Modern web applications are dynamic and interactive, featuring frequent user-page interactions. However, existing benchmarks largely evaluate the code generation of static webpages, ignoring the complex interactive behaviors in real-world applications. Besides, their evaluation criteria remain confined to visual fidelity and code structure, overlooking the interaction consistency between the generated and the reference webpages. To address these limitations, we introduce WebIGBench, the first benchmark designed to evaluate code generation for interactive webpages with complex interactions. By combining manually designed interaction paths with UI automation, we collected 103 complex webpages from real-world websites. This benchmark covers 5 popular interactive action types (e.g., click, input) involving 871 distinct interactive actions. Moreover, we propose a novel evaluation pipeline to address the gap in automated assessment of interactive actions. Extensive experiments on several representative MLLMs reveal the performance boundaries of current models in interactive webpage code generation using WebIGBench. The proposed benchmark is available at https://github.com/anoa12159-hue/WebIGBench_eval.
title Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2606.00154