Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
Fuente:
arXiv
Saved in:
| Main Authors: | He, Zehai, Hong, Wenyi, Yang, Zhen, Pan, Ziyang, Liu, Mingdao, Gu, Xiaotao, Tang, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WebVIA: A Web-based Vision-Language Agentic Framework for Interactive and Verifiable UI-to-Code Generation
by: Xu, Mingde, et al.
Published: (2025)
by: Xu, Mingde, et al.
Published: (2025)
The Kansei Engineering Approach in Web Design:Case of Transportation Website
by: Akram, Alisher, et al.
Published: (2024)
by: Akram, Alisher, et al.
Published: (2024)
StressWeb: A Diagnostic Benchmark for Web Agent Robustness under Realistic Interaction Variability
by: Bai, Haoyue, et al.
Published: (2026)
by: Bai, Haoyue, et al.
Published: (2026)
A Small Leak Sinks All: Exploring the Transferable Vulnerability of Source Code Models
by: Li, Weiye, et al.
Published: (2025)
by: Li, Weiye, et al.
Published: (2025)
Temac: Multi-Agent Collaboration for Automated Web GUI Testing
by: Liu, Chenxu, et al.
Published: (2025)
by: Liu, Chenxu, et al.
Published: (2025)
MACAA: Belief-Revision Multi-Agent Reasoning for Code Authorship Verification
by: Ye, Jingwei, et al.
Published: (2026)
by: Ye, Jingwei, et al.
Published: (2026)
Requirements Development and Formalization for Reliable Code Generation: A Multi-Agent Vision
by: Lu, Xu, et al.
Published: (2025)
by: Lu, Xu, et al.
Published: (2025)
Goedel-Code-Prover: Hierarchical Proof Search for Open State-of-the-Art Code Verification
by: Li, Zenan, et al.
Published: (2026)
by: Li, Zenan, et al.
Published: (2026)
LabUtopia: High-Fidelity Simulation and Hierarchical Benchmark for Scientific Embodied Agents
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
ProjectEval: A Benchmark for Programming Agents Automated Evaluation on Project-Level Code Generation
by: Liu, Kaiyuan, et al.
Published: (2025)
by: Liu, Kaiyuan, et al.
Published: (2025)
WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
by: Liu, Chenxu, et al.
Published: (2026)
by: Liu, Chenxu, et al.
Published: (2026)
WebApp1K: A Practical Code-Generation Benchmark for Web App Development
by: Cui, Yi
Published: (2024)
by: Cui, Yi
Published: (2024)
AI IDEs or Autonomous Agents? Measuring the Impact of Coding Agents on Software Development
by: Agarwal, Shyam, et al.
Published: (2026)
by: Agarwal, Shyam, et al.
Published: (2026)
Are Autonomous Web Agents Good Testers?
by: Chevrot, Antoine, et al.
Published: (2025)
by: Chevrot, Antoine, et al.
Published: (2025)
Benchmarking and Studying the LLM-based Agent System in End-to-End Software Development
by: Zeng, Zhengran, et al.
Published: (2025)
by: Zeng, Zhengran, et al.
Published: (2025)
HumanEvalComm: Benchmarking the Communication Competence of Code Generation for LLMs and LLM Agent
by: Wu, Jie JW, et al.
Published: (2024)
by: Wu, Jie JW, et al.
Published: (2024)
Agents4PLC: Automating Closed-loop PLC Code Generation and Verification in Industrial Control Systems using LLM-based Agents
by: Liu, Zihan, et al.
Published: (2024)
by: Liu, Zihan, et al.
Published: (2024)
HerAgent: Rethinking the Automated Environment Deployment via Hierarchical Test Pyramid
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
Deep Reinforcement Learning for Automated Web GUI Testing
by: Gu, Zhiyu, et al.
Published: (2025)
by: Gu, Zhiyu, et al.
Published: (2025)
A Hypergraph-based Formalization of Hierarchical Reactive Modules and a Compositional Verification Method
by: Ishii, Daisuke
Published: (2024)
by: Ishii, Daisuke
Published: (2024)
Leveraging Large Vision Language Model For Better Automatic Web GUI Testing
by: Wang, Siyi, et al.
Published: (2024)
by: Wang, Siyi, et al.
Published: (2024)
WebMAC: A Multi-Agent Collaborative Framework for Scenario Testing of Web Systems
by: Wan, Zhenyu, et al.
Published: (2026)
by: Wan, Zhenyu, et al.
Published: (2026)
EmbedAgent: Benchmarking Large Language Models in Embedded System Development
by: Xu, Ruiyang, et al.
Published: (2025)
by: Xu, Ruiyang, et al.
Published: (2025)
From Runnable to Shippable: Multi-Agent Test-Driven Development for Generating Full-Stack Web Applications from Requirements
by: Wan, Yuxuan, et al.
Published: (2026)
by: Wan, Yuxuan, et al.
Published: (2026)
A Benchmark for Evaluating Repository-Level Code Agents with Intermediate Reasoning on Feature Addition Task
by: Liu, Shuhan, et al.
Published: (2026)
by: Liu, Shuhan, et al.
Published: (2026)
Software Development Life Cycle Perspective: A Survey of Benchmarks for Code Large Language Models and Agents
by: Wang, Kaixin, et al.
Published: (2025)
by: Wang, Kaixin, et al.
Published: (2025)
VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents
by: Guo, JunJia, et al.
Published: (2026)
by: Guo, JunJia, et al.
Published: (2026)
AgentGuard: Runtime Verification of AI Agents
by: Koohestani, Roham
Published: (2025)
by: Koohestani, Roham
Published: (2025)
UniDebugger: Hierarchical Multi-Agent Framework for Unified Software Debugging
by: Lee, Cheryl, et al.
Published: (2024)
by: Lee, Cheryl, et al.
Published: (2024)
Issue Retrieval and Verification Enhanced Supplementary Code Comment Generation
by: Zou, Yanzhen, et al.
Published: (2025)
by: Zou, Yanzhen, et al.
Published: (2025)
LLM-Based Multi-Agent Systems for Software Engineering: Literature Review, Vision and the Road Ahead
by: He, Junda, et al.
Published: (2024)
by: He, Junda, et al.
Published: (2024)
RE-oriented Model Development with LLM Support and Deduction-based Verification
by: Klimek, Radoslaw
Published: (2025)
by: Klimek, Radoslaw
Published: (2025)
EmbeWebAgent: Embedding Web Agents into Any Customized UI
by: Ma, Chenyang, et al.
Published: (2026)
by: Ma, Chenyang, et al.
Published: (2026)
Code Review Agent Benchmark
by: Zhang, Yuntong, et al.
Published: (2026)
by: Zhang, Yuntong, et al.
Published: (2026)
ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
by: Lu, Pengrui, et al.
Published: (2026)
by: Lu, Pengrui, et al.
Published: (2026)
From Charts to Code: A Hierarchical Benchmark for Multimodal Models
by: Tang, Jiahao, et al.
Published: (2025)
by: Tang, Jiahao, et al.
Published: (2025)
A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback
by: Duan, Guoliang, et al.
Published: (2025)
by: Duan, Guoliang, et al.
Published: (2025)
Towards Speeding up Program Repair with Non-Autoregressive Model
by: Yang, Zhenyu, et al.
Published: (2025)
by: Yang, Zhenyu, et al.
Published: (2025)
APISENSOR: Robust Discovery of Web API from Runtime Traffic Logs
by: Yang, Yanjing, et al.
Published: (2026)
by: Yang, Yanjing, et al.
Published: (2026)
ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments
by: Ai, Jiaxin, et al.
Published: (2025)
by: Ai, Jiaxin, et al.
Published: (2025)
Similar Items
-
WebVIA: A Web-based Vision-Language Agentic Framework for Interactive and Verifiable UI-to-Code Generation
by: Xu, Mingde, et al.
Published: (2025) -
The Kansei Engineering Approach in Web Design:Case of Transportation Website
by: Akram, Alisher, et al.
Published: (2024) -
StressWeb: A Diagnostic Benchmark for Web Agent Robustness under Realistic Interaction Variability
by: Bai, Haoyue, et al.
Published: (2026) -
A Small Leak Sinks All: Exploring the Transferable Vulnerability of Source Code Models
by: Li, Weiye, et al.
Published: (2025) -
Temac: Multi-Agent Collaboration for Automated Web GUI Testing
by: Liu, Chenxu, et al.
Published: (2025)