WebTestBench: Evaluating Computer-Use Agents towards End-to-End Automated Web Testing
Fuente:
arXiv
Saved in:
| Main Authors: | Kong, Fanheng, Zhang, Jingyuan, Yue, Yang, Sun, Chenxi, Tian, Yang, Feng, Shi, Yang, Xiaocui, Wang, Daling, Tian, Yu, Du, Jun, Zeng, Wenchong, Li, Han, Gai, Kun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AgentWebBench: Benchmarking Multi-Agent Coordination in Agentic Web
by: Zhong, Shanshan, et al.
Published: (2026)
by: Zhong, Shanshan, et al.
Published: (2026)
FilmAgent: A Multi-Agent Framework for End-to-End Film Automation in Virtual 3D Spaces
by: Xu, Zhenran, et al.
Published: (2025)
by: Xu, Zhenran, et al.
Published: (2025)
An End-to-End Collaborative Learning Approach for Connected Autonomous Vehicles in Occluded Scenarios
by: Parada, Leandro, et al.
Published: (2024)
by: Parada, Leandro, et al.
Published: (2024)
Doctorina MedBench: End-to-End Evaluation of Agent-Based Medical AI
by: Kozlova, Anna, et al.
Published: (2026)
by: Kozlova, Anna, et al.
Published: (2026)
Dialogue Diplomats: An End-to-End Multi-Agent Reinforcement Learning System for Automated Conflict Resolution and Consensus Building
by: Bolleddu, Deepak
Published: (2025)
by: Bolleddu, Deepak
Published: (2025)
Holos: A Web-Scale LLM-Based Multi-Agent System for the Agentic Web
by: Nie, Xiaohang, et al.
Published: (2026)
by: Nie, Xiaohang, et al.
Published: (2026)
The AI Committee: A Multi-Agent Framework for Automated Validation and Remediation of Web-Sourced Data
by: Vallabhaneni, Sunith, et al.
Published: (2025)
by: Vallabhaneni, Sunith, et al.
Published: (2025)
SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies
by: Saxena, Siddhant, et al.
Published: (2026)
by: Saxena, Siddhant, et al.
Published: (2026)
End-to-End Autonomous Driving through V2X Cooperation
by: Yu, Haibao, et al.
Published: (2024)
by: Yu, Haibao, et al.
Published: (2024)
Semantic Web Technology for Agent Communication Protocols
by: Berges, Idoia, et al.
Published: (2024)
by: Berges, Idoia, et al.
Published: (2024)
StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning
by: Li, Shiyang, et al.
Published: (2026)
by: Li, Shiyang, et al.
Published: (2026)
What if Pinocchio Were a Reinforcement Learning Agent: A Normative End-to-End Pipeline
by: Alcaraz, Benoît
Published: (2026)
by: Alcaraz, Benoît
Published: (2026)
Unified End-to-End V2X Cooperative Autonomous Driving
by: Li, Zhiwei, et al.
Published: (2024)
by: Li, Zhiwei, et al.
Published: (2024)
Learning to Communicate: Toward End-to-End Optimization of Multi-Agent Language Systems
by: Yu, Ye, et al.
Published: (2026)
by: Yu, Ye, et al.
Published: (2026)
LiteWebAgent: The Open-Source Suite for VLM-Based Web-Agent Applications
by: Zhang, Danqing, et al.
Published: (2025)
by: Zhang, Danqing, et al.
Published: (2025)
Synergy: A Next-Generation General-Purpose Agent for Open Agentic Web
by: Nie, Xiaohang, et al.
Published: (2026)
by: Nie, Xiaohang, et al.
Published: (2026)
Beyond Browsing: API-Based Web Agents
by: Song, Yueqi, et al.
Published: (2024)
by: Song, Yueqi, et al.
Published: (2024)
UCAgent: An End-to-End Agent for Block-Level Functional Verification
by: Wang, Junyue, et al.
Published: (2026)
by: Wang, Junyue, et al.
Published: (2026)
Agents on the Bench: Large Language Model Based Multi Agent Framework for Trustworthy Digital Justice
by: Jiang, Cong, et al.
Published: (2024)
by: Jiang, Cong, et al.
Published: (2024)
It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents
by: Korgul, Karolina, et al.
Published: (2025)
by: Korgul, Karolina, et al.
Published: (2025)
Manalyzer: End-to-end Automated Meta-analysis with Multi-agent System
by: Xu, Wanghan, et al.
Published: (2025)
by: Xu, Wanghan, et al.
Published: (2025)
MedPriv-Bench: Benchmarking the Privacy-Utility Trade-off of Large Language Models in Medical Open-End Question Answering
by: Guan, Shaowei, et al.
Published: (2026)
by: Guan, Shaowei, et al.
Published: (2026)
Web Fraud Attacks Against LLM-Driven Multi-Agent Systems
by: Kong, Dezhang, et al.
Published: (2025)
by: Kong, Dezhang, et al.
Published: (2025)
MLZero: A Multi-Agent System for End-to-end Machine Learning Automation
by: Fang, Haoyang, et al.
Published: (2025)
by: Fang, Haoyang, et al.
Published: (2025)
From Grounding to Planning: Benchmarking Bottlenecks in Web Agents
by: Shlomov, Segev, et al.
Published: (2024)
by: Shlomov, Segev, et al.
Published: (2024)
Cognitive Duality for Adaptive Web Agents
by: Liu, Jiarun, et al.
Published: (2025)
by: Liu, Jiarun, et al.
Published: (2025)
A Comprehensive Survey on Multi-Agent Cooperative Decision-Making: Scenarios, Approaches, Challenges and Perspectives
by: Jin, Weiqiang, et al.
Published: (2025)
by: Jin, Weiqiang, et al.
Published: (2025)
A BDI Agent-Based Task Scheduling Framework for Cloud Computing
by: Yang, Yikun, et al.
Published: (2024)
by: Yang, Yikun, et al.
Published: (2024)
SCP: Accelerating Discovery with a Global Web of Autonomous Scientific Agents
by: Jiang, Yankai, et al.
Published: (2025)
by: Jiang, Yankai, et al.
Published: (2025)
Prune4Web: DOM Tree Pruning Programming for Web Agent
by: Zhang, Jiayuan, et al.
Published: (2025)
by: Zhang, Jiayuan, et al.
Published: (2025)
LEARN: Learning End-to-End Aerial Resource-Constrained Multi-Robot Navigation
by: Chiu, Darren, et al.
Published: (2025)
by: Chiu, Darren, et al.
Published: (2025)
Symbiotic Cooperation for Web Agents: Harnessing Complementary Strengths of Large and Small LLMs
by: Zhang, Ruichen, et al.
Published: (2025)
by: Zhang, Ruichen, et al.
Published: (2025)
Agent-Oriented Visual Programming for the Web of Things
by: Burattini, Samuele, et al.
Published: (2025)
by: Burattini, Samuele, et al.
Published: (2025)
MAS-on-the-Fly: Dynamic Adaptation of LLM-based Multi-Agent Systems at Test Time
by: Liu, Guangyi, et al.
Published: (2026)
by: Liu, Guangyi, et al.
Published: (2026)
Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models
by: Choi, Younwoo, et al.
Published: (2025)
by: Choi, Younwoo, et al.
Published: (2025)
Agentic Automation of BT-RADS Scoring: End-to-End Multi-Agent System for Standardized Brain Tumor Follow-up Assessment
by: Jabal, Mohamed Sobhi, et al.
Published: (2026)
by: Jabal, Mohamed Sobhi, et al.
Published: (2026)
Bottom-Up Reputation Promotes Cooperation with Multi-Agent Reinforcement Learning
by: Ren, Tianyu, et al.
Published: (2025)
by: Ren, Tianyu, et al.
Published: (2025)
Collision Avoidance and Navigation for a Quadrotor Swarm Using End-to-end Deep Reinforcement Learning
by: Huang, Zhehui, et al.
Published: (2023)
by: Huang, Zhehui, et al.
Published: (2023)
MAPPO-LCR: Multi-Agent Proximal Policy Optimization with Local Cooperation Reward in Spatial Public Goods Games
by: Yang, Zhaoqilin, et al.
Published: (2025)
by: Yang, Zhaoqilin, et al.
Published: (2025)
Silo-Bench: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems
by: Zhang, Yuzhe, et al.
Published: (2026)
by: Zhang, Yuzhe, et al.
Published: (2026)
Similar Items
-
AgentWebBench: Benchmarking Multi-Agent Coordination in Agentic Web
by: Zhong, Shanshan, et al.
Published: (2026) -
FilmAgent: A Multi-Agent Framework for End-to-End Film Automation in Virtual 3D Spaces
by: Xu, Zhenran, et al.
Published: (2025) -
An End-to-End Collaborative Learning Approach for Connected Autonomous Vehicles in Occluded Scenarios
by: Parada, Leandro, et al.
Published: (2024) -
Doctorina MedBench: End-to-End Evaluation of Agent-Based Medical AI
by: Kozlova, Anna, et al.
Published: (2026) -
Dialogue Diplomats: An End-to-End Multi-Agent Reinforcement Learning System for Automated Conflict Resolution and Consensus Building
by: Bolleddu, Deepak
Published: (2025)