VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, JunJia, Yao, Yuhang, Jiawei, Zhou, Chen, Jingdi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WebCode2M: A Real-World Dataset for Code Generation from Webpage Designs
by: Gui, Yi, et al.
Published: (2024)
by: Gui, Yi, et al.
Published: (2024)
FronTalk: Benchmarking Front-End Development as Conversational Code Generation with Multi-Modal Feedback
by: Wu, Xueqing, et al.
Published: (2025)
by: Wu, Xueqing, et al.
Published: (2025)
Detecting and Characterising Mobile App Metamorphosis in Google Play Store
by: Denipitiyage, D., et al.
Published: (2024)
by: Denipitiyage, D., et al.
Published: (2024)
ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
by: Lu, Pengrui, et al.
Published: (2026)
by: Lu, Pengrui, et al.
Published: (2026)
JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligence
by: Sun, Qiushi, et al.
Published: (2025)
by: Sun, Qiushi, et al.
Published: (2025)
SVRepair: Structured Visual Reasoning for Automated Program Repair
by: Tang, Xiaoxuan, et al.
Published: (2026)
by: Tang, Xiaoxuan, et al.
Published: (2026)
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
by: Meng, Fanqing, et al.
Published: (2026)
by: Meng, Fanqing, et al.
Published: (2026)
WebApp1K: A Practical Code-Generation Benchmark for Web App Development
by: Cui, Yi
Published: (2024)
by: Cui, Yi
Published: (2024)
BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD
by: Zhang, Haozhe, et al.
Published: (2026)
by: Zhang, Haozhe, et al.
Published: (2026)
Can Vision-Language Models Handle Long-Context Code? An Empirical Study on Visual Compression
by: Zhong, Jianping, et al.
Published: (2026)
by: Zhong, Jianping, et al.
Published: (2026)
WAFFLE: Finetuning Multi-Modal Models for Automated Front-End Development
by: Liang, Shanchao, et al.
Published: (2024)
by: Liang, Shanchao, et al.
Published: (2024)
CUARewardBench: A Benchmark for Evaluating Reward Models on Computer-using Agent
by: Lin, Haojia, et al.
Published: (2025)
by: Lin, Haojia, et al.
Published: (2025)
MMCode: Benchmarking Multimodal Large Language Models for Code Generation with Visually Rich Programming Problems
by: Li, Kaixin, et al.
Published: (2024)
by: Li, Kaixin, et al.
Published: (2024)
Extracting Overlapping Microservices from Monolithic Code via Deep Semantic Embeddings and Graph Neural Network-Based Soft Clustering
by: Ziabakhsh, Morteza, et al.
Published: (2025)
by: Ziabakhsh, Morteza, et al.
Published: (2025)
Revisiting Out-of-Distribution Detection in Real-time Object Detection: From Benchmark Pitfalls to a New Mitigation Paradigm
by: Wu, Changshun, et al.
Published: (2025)
by: Wu, Changshun, et al.
Published: (2025)
FullStack-Agent: Enhancing Agentic Full-Stack Web Coding via Development-Oriented Testing and Repository Back-Translation
by: Lu, Zimu, et al.
Published: (2026)
by: Lu, Zimu, et al.
Published: (2026)
DOne: Decoupling Structure and Rendering for High-Fidelity Design-to-Code Generation
by: Huang, Xinhao, et al.
Published: (2026)
by: Huang, Xinhao, et al.
Published: (2026)
Insights from Benchmarking Frontier Language Models on Web App Code Generation
by: Cui, Yi
Published: (2024)
by: Cui, Yi
Published: (2024)
Computer Vision Intelligence Test Modeling and Generation: A Case Study on Smart OCR
by: Shu, Jing, et al.
Published: (2024)
by: Shu, Jing, et al.
Published: (2024)
GPA: Learning GUI Process Automation from Demonstrations
by: Zhao, Zirui, et al.
Published: (2026)
by: Zhao, Zirui, et al.
Published: (2026)
SpecAgent: A Speculative Retrieval and Forecasting Agent for Code Completion
by: Ma, George, et al.
Published: (2025)
by: Ma, George, et al.
Published: (2025)
GUI Agents for Continual Game Generation
by: Huang, Yixu, et al.
Published: (2026)
by: Huang, Yixu, et al.
Published: (2026)
Interpretable Gallbladder Ultrasound Diagnosis: A Lightweight Web-Mobile Software Platform with Real-Time XAI
by: Bhoyan, Fuyad Hasan, et al.
Published: (2025)
by: Bhoyan, Fuyad Hasan, et al.
Published: (2025)
WebTestBench: Evaluating Computer-Use Agents towards End-to-End Automated Web Testing
by: Kong, Fanheng, et al.
Published: (2026)
by: Kong, Fanheng, et al.
Published: (2026)
Benchmarking Image Perturbations for Testing Automated Driving Assistance Systems
by: Lambertenghi, Stefano Carlo, et al.
Published: (2025)
by: Lambertenghi, Stefano Carlo, et al.
Published: (2025)
Agents in the Sandbox: End-to-End Crash Bug Reproduction for Minecraft
by: Yapağcı, Eray, et al.
Published: (2025)
by: Yapağcı, Eray, et al.
Published: (2025)
How Smart Is Your GUI Agent? A Framework for the Future of Software Interaction
by: Feng, Sidong, et al.
Published: (2026)
by: Feng, Sidong, et al.
Published: (2026)
VEglue: Testing Visual Entailment Systems via Object-Aligned Joint Erasing
by: Chang, Zhiyuan, et al.
Published: (2024)
by: Chang, Zhiyuan, et al.
Published: (2024)
EdgeFlow: Edge-Map Augmented VLM-Based Flowchart Processing for Industrial Requirements Engineering
by: Dou, Zhifei, et al.
Published: (2026)
by: Dou, Zhifei, et al.
Published: (2026)
Metamorphic Testing for Pose Estimation Systems
by: Duran, Matias, et al.
Published: (2025)
by: Duran, Matias, et al.
Published: (2025)
Training-Free Consistency Pipeline for Fashion Repose
by: Aghilar, Potito, et al.
Published: (2025)
by: Aghilar, Potito, et al.
Published: (2025)
An Architecture-Led Hybrid Report on Body Language Detection Project
by: Tong, Thomson, et al.
Published: (2025)
by: Tong, Thomson, et al.
Published: (2025)
EvoDev: An Iterative Feature-Driven Framework for End-to-End Software Development with LLM-based Agents
by: Liu, Junwei, et al.
Published: (2025)
by: Liu, Junwei, et al.
Published: (2025)
Runtime Failure Hunting for Physics Engine Based Software Systems: How Far Can We Go?
by: Li, Shuqing, et al.
Published: (2025)
by: Li, Shuqing, et al.
Published: (2025)
Tree-of-Code: A Tree-Structured Exploring Framework for End-to-End Code Generation and Execution in Complex Task Handling
by: Ni, Ziyi, et al.
Published: (2024)
by: Ni, Ziyi, et al.
Published: (2024)
How Far Can VLMs Go for Visual Bug Detection? Studying 19,738 Keyframes from 41 Hours of Gameplay Videos
by: Lu, Wentao, et al.
Published: (2026)
by: Lu, Wentao, et al.
Published: (2026)
How Well Do Large Language Models Serve as End-to-End Secure Code Agents for Python?
by: Gong, Jianian, et al.
Published: (2024)
by: Gong, Jianian, et al.
Published: (2024)
Natural Adversaries: Fuzzing Autonomous Vehicles with Realistic Roadside Object Placements
by: Sun, Yang, et al.
Published: (2024)
by: Sun, Yang, et al.
Published: (2024)
MotorEase: Automated Detection of Motor Impairment Accessibility Issues in Mobile App UIs
by: Krishnavajjala, Arun, et al.
Published: (2024)
by: Krishnavajjala, Arun, et al.
Published: (2024)
CodeFuse-CR-Bench: A Comprehensiveness-aware Benchmark for End-to-End Code Review Evaluation in Python Projects
by: Guo, Hanyang, et al.
Published: (2025)
by: Guo, Hanyang, et al.
Published: (2025)
Similar Items
-
WebCode2M: A Real-World Dataset for Code Generation from Webpage Designs
by: Gui, Yi, et al.
Published: (2024) -
FronTalk: Benchmarking Front-End Development as Conversational Code Generation with Multi-Modal Feedback
by: Wu, Xueqing, et al.
Published: (2025) -
Detecting and Characterising Mobile App Metamorphosis in Google Play Store
by: Denipitiyage, D., et al.
Published: (2024) -
ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
by: Lu, Pengrui, et al.
Published: (2026) -
JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligence
by: Sun, Qiushi, et al.
Published: (2025)