InteractWeb-Bench: Can Multimodal Agent Escape Blind Execution in Interactive Website Generation?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Qiyao, Hu, Haoran, Chen, Longze, Wang, Hongbo, Alinejad-Rokny, Hamid, Lin, Yuan, Yang, Min
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910179621076992
author Wang, Qiyao
Hu, Haoran
Chen, Longze
Wang, Hongbo
Alinejad-Rokny, Hamid
Lin, Yuan
Yang, Min
author_facet Wang, Qiyao
Hu, Haoran
Chen, Longze
Wang, Hongbo
Alinejad-Rokny, Hamid
Lin, Yuan
Yang, Min
contents With the advancement of multimodal large language models (MLLMs) and coding agents, the website development has shifted from manual programming to agent-based project-level code synthesis. Existing benchmarks rely on idealized assumptions, especially for well-structured, information-rich inputs and static execution settings. In contrast, real-world development is constrained by a critical bottleneck: the semantic misalignment between ambiguous, low-quality instructions from non-expert users and model understanding, which results in a failure mode that we term blind execution. To address this gap, we introduce InteractWeb-Bench, the first multimodal interactive benchmark for website generation under non-expert low-code user conditions. InteractWeb-Bench introduces four types of user agents and persona-driven instruction perturbations to systematically simulate diverse user behaviors, including ambiguity, redundancy, and contradiction, grounded in requirement engineering defect taxonomies. We develop an interactive execution environment for agents, featuring a unified action space comprising Clarify, Implement, Verify, and Submit, enabling iterative intent refinement, code synthesis, and visual feedback-based validation. Extensive experiments and analysis reveal that frontier MLLM-based agents remain trapped in blind execution, exposing limitations in intent recognition and adaptive interaction.
format Preprint
id arxiv_https___arxiv_org_abs_2604_27419
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle InteractWeb-Bench: Can Multimodal Agent Escape Blind Execution in Interactive Website Generation?
Wang, Qiyao
Hu, Haoran
Chen, Longze
Wang, Hongbo
Alinejad-Rokny, Hamid
Lin, Yuan
Yang, Min
Artificial Intelligence
Computation and Language
With the advancement of multimodal large language models (MLLMs) and coding agents, the website development has shifted from manual programming to agent-based project-level code synthesis. Existing benchmarks rely on idealized assumptions, especially for well-structured, information-rich inputs and static execution settings. In contrast, real-world development is constrained by a critical bottleneck: the semantic misalignment between ambiguous, low-quality instructions from non-expert users and model understanding, which results in a failure mode that we term blind execution. To address this gap, we introduce InteractWeb-Bench, the first multimodal interactive benchmark for website generation under non-expert low-code user conditions. InteractWeb-Bench introduces four types of user agents and persona-driven instruction perturbations to systematically simulate diverse user behaviors, including ambiguity, redundancy, and contradiction, grounded in requirement engineering defect taxonomies. We develop an interactive execution environment for agents, featuring a unified action space comprising Clarify, Implement, Verify, and Submit, enabling iterative intent refinement, code synthesis, and visual feedback-based validation. Extensive experiments and analysis reveal that frontier MLLM-based agents remain trapped in blind execution, exposing limitations in intent recognition and adaptive interaction.
title InteractWeb-Bench: Can Multimodal Agent Escape Blind Execution in Interactive Website Generation?
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2604.27419