StressWeb: A Diagnostic Benchmark for Web Agent Robustness under Realistic Interaction Variability
Fuente:
arXiv
Guardado en:
| Autores principales: | Bai, Haoyue, Wang, Dong, Chen, Long, Hao, Bingguang, Shao, Pengyang, Yang, Yonghui, He, Yicheng, Zhuang, Chenyi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
EmbeWebAgent: Embedding Web Agents into Any Customized UI
por: Ma, Chenyang, et al.
Publicado: (2026)
por: Ma, Chenyang, et al.
Publicado: (2026)
Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
por: He, Zehai, et al.
Publicado: (2026)
por: He, Zehai, et al.
Publicado: (2026)
WebSuite: Systematically Evaluating Why Web Agents Fail
por: Li, Eric, et al.
Publicado: (2024)
por: Li, Eric, et al.
Publicado: (2024)
Are Autonomous Web Agents Good Testers?
por: Chevrot, Antoine, et al.
Publicado: (2025)
por: Chevrot, Antoine, et al.
Publicado: (2025)
WebMAC: A Multi-Agent Collaborative Framework for Scenario Testing of Web Systems
por: Wan, Zhenyu, et al.
Publicado: (2026)
por: Wan, Zhenyu, et al.
Publicado: (2026)
APISENSOR: Robust Discovery of Web API from Runtime Traffic Logs
por: Yang, Yanjing, et al.
Publicado: (2026)
por: Yang, Yanjing, et al.
Publicado: (2026)
WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
por: Liu, Chenxu, et al.
Publicado: (2026)
por: Liu, Chenxu, et al.
Publicado: (2026)
WebApp1K: A Practical Code-Generation Benchmark for Web App Development
por: Cui, Yi
Publicado: (2024)
por: Cui, Yi
Publicado: (2024)
Temac: Multi-Agent Collaboration for Automated Web GUI Testing
por: Liu, Chenxu, et al.
Publicado: (2025)
por: Liu, Chenxu, et al.
Publicado: (2025)
Debugging WebAssembly? Put some Whamm on it!
por: Gilbert, Elizabeth, et al.
Publicado: (2025)
por: Gilbert, Elizabeth, et al.
Publicado: (2025)
ContractSkill: Repairable Contract-Based Skills for Multimodal Web Agents
por: Lu, Zijian, et al.
Publicado: (2026)
por: Lu, Zijian, et al.
Publicado: (2026)
WebSPL: A Software Product Line for Web Applications
por: da Luz, Maicon Azevedo, et al.
Publicado: (2024)
por: da Luz, Maicon Azevedo, et al.
Publicado: (2024)
Envisioning Future Interactive Web Development: Editing Webpage with Natural Language
por: Dang, Truong Hai, et al.
Publicado: (2025)
por: Dang, Truong Hai, et al.
Publicado: (2025)
WebVIA: A Web-based Vision-Language Agentic Framework for Interactive and Verifiable UI-to-Code Generation
por: Xu, Mingde, et al.
Publicado: (2025)
por: Xu, Mingde, et al.
Publicado: (2025)
Multimodal Auto Validation For Self-Refinement in Web Agents
por: Azam, Ruhana, et al.
Publicado: (2024)
por: Azam, Ruhana, et al.
Publicado: (2024)
AutoWebWorld: Synthesizing Infinite Verifiable Web Environments via Finite State Machines
por: Wu, Yifan, et al.
Publicado: (2026)
por: Wu, Yifan, et al.
Publicado: (2026)
VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents
por: Guo, JunJia, et al.
Publicado: (2026)
por: Guo, JunJia, et al.
Publicado: (2026)
WebTestPilot: Agentic End-to-End Web Testing against Natural Language Specification by Inferring Oracles with Symbolized GUI Elements
por: Teoh, Xiwen, et al.
Publicado: (2026)
por: Teoh, Xiwen, et al.
Publicado: (2026)
The Promise and Pitfalls of WebAssembly: Perspectives from the Industry
por: He, Ningyu, et al.
Publicado: (2025)
por: He, Ningyu, et al.
Publicado: (2025)
Autonomous Legacy Web Application Upgrades Using a Multi-Agent System
por: Ala-Salmi, Valtteri, et al.
Publicado: (2025)
por: Ala-Salmi, Valtteri, et al.
Publicado: (2025)
DSCodeBench: A Realistic Benchmark for Data Science Code Generation
por: Ouyang, Shuyin, et al.
Publicado: (2025)
por: Ouyang, Shuyin, et al.
Publicado: (2025)
Web Element Relocalization in Evolving Web Applications: A Comparative Analysis and Extension Study
por: Kluge, Anton, et al.
Publicado: (2025)
por: Kluge, Anton, et al.
Publicado: (2025)
WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models
por: Lei, Xinping, et al.
Publicado: (2026)
por: Lei, Xinping, et al.
Publicado: (2026)
The BrowserGym Ecosystem for Web Agent Research
por: De Chezelles, Thibault Le Sellier, et al.
Publicado: (2024)
por: De Chezelles, Thibault Le Sellier, et al.
Publicado: (2024)
LLM Agents for Automated Web Vulnerability Reproduction: Are We There Yet?
por: Liu, Bin, et al.
Publicado: (2025)
por: Liu, Bin, et al.
Publicado: (2025)
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
por: Merrill, Mike A., et al.
Publicado: (2026)
por: Merrill, Mike A., et al.
Publicado: (2026)
Neural Embeddings for Web Testing
por: Kanaththage, Kasun, et al.
Publicado: (2023)
por: Kanaththage, Kasun, et al.
Publicado: (2023)
An Autonomous RL Agent Methodology for Dynamic Web UI Testing in a BDD Framework
por: Mughal, Ali Hassaan
Publicado: (2025)
por: Mughal, Ali Hassaan
Publicado: (2025)
Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation
por: Garg, Spandan, et al.
Publicado: (2025)
por: Garg, Spandan, et al.
Publicado: (2025)
Agent-Oriented Visual Programming for the Web of Things
por: Burattini, Samuele, et al.
Publicado: (2025)
por: Burattini, Samuele, et al.
Publicado: (2025)
Development of an Automated Web Application for Efficient Web Scraping: Design and Implementation
por: Dutta, Alok, et al.
Publicado: (2025)
por: Dutta, Alok, et al.
Publicado: (2025)
CodeCircuit: Toward Inferring LLM-Generated Code Correctness via Attribution Graphs
por: He, Yicheng, et al.
Publicado: (2026)
por: He, Yicheng, et al.
Publicado: (2026)
Insights from Benchmarking Frontier Language Models on Web App Code Generation
por: Cui, Yi
Publicado: (2024)
por: Cui, Yi
Publicado: (2024)
WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality
por: Li, Chunyang, et al.
Publicado: (2025)
por: Li, Chunyang, et al.
Publicado: (2025)
WebTestBench: Evaluating Computer-Use Agents towards End-to-End Automated Web Testing
por: Kong, Fanheng, et al.
Publicado: (2026)
por: Kong, Fanheng, et al.
Publicado: (2026)
Accessibility Issues in Ad-Driven Web Applications
por: Amjad, Abdul Haddi, et al.
Publicado: (2024)
por: Amjad, Abdul Haddi, et al.
Publicado: (2024)
Testing Medical Rules Web Services in Practice
por: Laaber, Christoph, et al.
Publicado: (2024)
por: Laaber, Christoph, et al.
Publicado: (2024)
Research on WebAssembly Runtimes: A Survey
por: Zhang, Yixuan, et al.
Publicado: (2024)
por: Zhang, Yixuan, et al.
Publicado: (2024)
Do RESTful API Design Rules Have an Impact on the Understandability of Web APIs? A Web-Based Experiment with API Descriptions
por: Bogner, Justus, et al.
Publicado: (2023)
por: Bogner, Justus, et al.
Publicado: (2023)
Energy Patterns for Web: An Exploratory Study
por: Rani, Pooja, et al.
Publicado: (2024)
por: Rani, Pooja, et al.
Publicado: (2024)
Ejemplares similares
-
EmbeWebAgent: Embedding Web Agents into Any Customized UI
por: Ma, Chenyang, et al.
Publicado: (2026) -
Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
por: He, Zehai, et al.
Publicado: (2026) -
WebSuite: Systematically Evaluating Why Web Agents Fail
por: Li, Eric, et al.
Publicado: (2024) -
Are Autonomous Web Agents Good Testers?
por: Chevrot, Antoine, et al.
Publicado: (2025) -
WebMAC: A Multi-Agent Collaborative Framework for Scenario Testing of Web Systems
por: Wan, Zhenyu, et al.
Publicado: (2026)