Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Zehai, Hong, Wenyi, Yang, Zhen, Pan, Ziyang, Liu, Mingdao, Gu, Xiaotao, Tang, Jie
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908930071855104
author He, Zehai
Hong, Wenyi
Yang, Zhen
Pan, Ziyang
Liu, Mingdao
Gu, Xiaotao
Tang, Jie
author_facet He, Zehai
Hong, Wenyi
Yang, Zhen
Pan, Ziyang
Liu, Mingdao
Gu, Xiaotao
Tang, Jie
contents Recent advances in large language models have improved the capabilities of coding agents, yet systematic evaluation of complex, end-to-end website development remains limited. To address this gap, we introduce Vision2Web, a hierarchical benchmark for visual website development, spanning from static UI-to-code generation, interactive multi-page frontend reproduction, to long-horizon full-stack website development. The benchmark is constructed from real-world websites and comprises a total of 193 tasks across 16 categories, with 918 prototype images and 1,255 test cases. To support flexible, thorough and reliable evaluation, we propose workflow-based agent verification paradigm based on two complementary components: a GUI agent verifier and a VLM-based judge. We evaluate multiple visual language models instantiated under different coding-agent frameworks, revealing substantial performance gaps at all task levels, with state-of-the-art models still struggling on full-stack development.
format Preprint
id arxiv_https___arxiv_org_abs_2603_26648
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
He, Zehai
Hong, Wenyi
Yang, Zhen
Pan, Ziyang
Liu, Mingdao
Gu, Xiaotao
Tang, Jie
Software Engineering
Artificial Intelligence
Recent advances in large language models have improved the capabilities of coding agents, yet systematic evaluation of complex, end-to-end website development remains limited. To address this gap, we introduce Vision2Web, a hierarchical benchmark for visual website development, spanning from static UI-to-code generation, interactive multi-page frontend reproduction, to long-horizon full-stack website development. The benchmark is constructed from real-world websites and comprises a total of 193 tasks across 16 categories, with 918 prototype images and 1,255 test cases. To support flexible, thorough and reliable evaluation, we propose workflow-based agent verification paradigm based on two complementary components: a GUI agent verifier and a VLM-based judge. We evaluate multiple visual language models instantiated under different coding-agent frameworks, revealing substantial performance gaps at all task levels, with state-of-the-art models still struggling on full-stack development.
title Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2603.26648