Training Computer Use Agents to Assess the Usability of Graphical User Interfaces
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Alice, Tong, Weixi, Vempati, Rishab, Reinecke, Katharina, Shapiro, R. Benjamin, Zhang, Tianyi, Wu, Jason |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mango: Multi-Agent Web Navigation via Global-View Optimization
by: Tong, Weixi, et al.
Published: (2026)
by: Tong, Weixi, et al.
Published: (2026)
GraphicBench: A Planning Benchmark for Graphic Design with Language Agents
by: Ki, Dayeon, et al.
Published: (2025)
by: Ki, Dayeon, et al.
Published: (2025)
Efficient Agent Training for Computer Use
by: He, Yanheng, et al.
Published: (2025)
by: He, Yanheng, et al.
Published: (2025)
An Efficient Training Pipeline for Reasoning Graphical User Interface Agents
by: Pantazopoulos, Georgios, et al.
Published: (2025)
by: Pantazopoulos, Georgios, et al.
Published: (2025)
CodeJudge: Evaluating Code Generation with Large Language Models
by: Tong, Weixi, et al.
Published: (2024)
by: Tong, Weixi, et al.
Published: (2024)
Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis
by: Xie, Tianbao, et al.
Published: (2025)
by: Xie, Tianbao, et al.
Published: (2025)
Computer-Use Agents as Judges for Generative User Interface
by: Lin, Kevin Qinghong, et al.
Published: (2025)
by: Lin, Kevin Qinghong, et al.
Published: (2025)
On the Suitability of LLM-Driven Agents for Dark Pattern Audits
by: Sun, Chen, et al.
Published: (2026)
by: Sun, Chen, et al.
Published: (2026)
LM Agents for Coordinating Multi-User Information Gathering
by: Jhamtani, Harsh, et al.
Published: (2025)
by: Jhamtani, Harsh, et al.
Published: (2025)
Human-Guided Harm Recovery for Computer Use Agents
by: Li, Christy, et al.
Published: (2026)
by: Li, Christy, et al.
Published: (2026)
AutoGRAMS: Autonomous Graphical Agent Modeling Software
by: Krause, Ben, et al.
Published: (2024)
by: Krause, Ben, et al.
Published: (2024)
UserRL: Training Interactive User-Centric Agent via Reinforcement Learning
by: Qian, Cheng, et al.
Published: (2025)
by: Qian, Cheng, et al.
Published: (2025)
Creating General User Models from Computer Use
by: Shaikh, Omar, et al.
Published: (2025)
by: Shaikh, Omar, et al.
Published: (2025)
The Fire Thief Is Also the Keeper: Balancing Usability and Privacy in Prompts
by: Shen, Zhili, et al.
Published: (2024)
by: Shen, Zhili, et al.
Published: (2024)
An Automatic Question Usability Evaluation Toolkit
by: Moore, Steven, et al.
Published: (2024)
by: Moore, Steven, et al.
Published: (2024)
Enhancing Dialogue State Tracking Models through LLM-backed User-Agents Simulation
by: Niu, Cheng, et al.
Published: (2024)
by: Niu, Cheng, et al.
Published: (2024)
VideoAgentTrek: Computer Use Pretraining from Unlabeled Videos
by: Lu, Dunjie, et al.
Published: (2025)
by: Lu, Dunjie, et al.
Published: (2025)
Scaling Agents for Computer Use
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
Generation Z's Ability to Discriminate Between AI-generated and Human-Authored Text on Discord
by: Ramu, Dhruv, et al.
Published: (2023)
by: Ramu, Dhruv, et al.
Published: (2023)
Interpreting User Requests in the Context of Natural Language Standing Instructions
by: Moghe, Nikita, et al.
Published: (2023)
by: Moghe, Nikita, et al.
Published: (2023)
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
by: Sun, Zeyi, et al.
Published: (2025)
by: Sun, Zeyi, et al.
Published: (2025)
Structured Reasoning for Fairness: A Multi-Agent Approach to Bias Detection in Textual Data
by: Huang, Tianyi, et al.
Published: (2025)
by: Huang, Tianyi, et al.
Published: (2025)
The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use
by: Hu, Siyuan, et al.
Published: (2024)
by: Hu, Siyuan, et al.
Published: (2024)
Low-code LLM: Graphical User Interface over Large Language Models
by: Cai, Yuzhe, et al.
Published: (2023)
by: Cai, Yuzhe, et al.
Published: (2023)
Higress-RAG: A Holistic Optimization Framework for Enterprise Retrieval-Augmented Generation via Dual Hybrid Retrieval, Adaptive Routing, and CRAG
by: Lin, Weixi
Published: (2025)
by: Lin, Weixi
Published: (2025)
A Graph Talks, But Who's Listening? Rethinking Evaluations for Graph-Language Models
by: Petkar, Soham, et al.
Published: (2025)
by: Petkar, Soham, et al.
Published: (2025)
MedExAgent: Training LLM Agents to Ask, Examine, and Diagnose in Noisy Clinical Environments
by: Gao, Yicheng, et al.
Published: (2026)
by: Gao, Yicheng, et al.
Published: (2026)
UserHarness: Harnessing User Minds for Stronger Agent Theory-of-Mind
by: Qian, Cheng, et al.
Published: (2026)
by: Qian, Cheng, et al.
Published: (2026)
Beyond Single-User Dialogue: Assessing Multi-User Dialogue State Tracking Capabilities of Large Language Models
by: Song, Sangmin, et al.
Published: (2025)
by: Song, Sangmin, et al.
Published: (2025)
UIClip: A Data-driven Model for Assessing User Interface Design
by: Wu, Jason, et al.
Published: (2024)
by: Wu, Jason, et al.
Published: (2024)
DiagrammaticLearning: A Graphical Language for Compositional Training Regimes
by: Lary, Mason, et al.
Published: (2025)
by: Lary, Mason, et al.
Published: (2025)
GUIDE: Graphical User Interface Data for Execution
by: Chawla, Rajat, et al.
Published: (2024)
by: Chawla, Rajat, et al.
Published: (2024)
Testing and Understanding Erroneous Planning in LLM Agents through Synthesized User Inputs
by: Ji, Zhenlan, et al.
Published: (2024)
by: Ji, Zhenlan, et al.
Published: (2024)
VehicleMemBench: An Executable Benchmark for Multi-User Long-Term Memory in In-Vehicle Agents
by: Chen, Yuhao, et al.
Published: (2026)
by: Chen, Yuhao, et al.
Published: (2026)
Learn from Weaknesses: Automated Domain Specialization for Small Computer-Use Agents
by: Kim, Suji, et al.
Published: (2026)
by: Kim, Suji, et al.
Published: (2026)
GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL
by: Yang, Rui, et al.
Published: (2026)
by: Yang, Rui, et al.
Published: (2026)
OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
by: Hu, Xueyu, et al.
Published: (2025)
by: Hu, Xueyu, et al.
Published: (2025)
From Self-Evolving Synthetic Data to Verifiable-Reward RL: Post-Training Multi-turn Interactive Tool-Using Agents
by: Gao, Jiaxuan, et al.
Published: (2026)
by: Gao, Jiaxuan, et al.
Published: (2026)
Toward Metaphor-Fluid Conversation Design for Voice User Interfaces
by: Desai, Smit, et al.
Published: (2025)
by: Desai, Smit, et al.
Published: (2025)
Memory-Augmented Agent Training for Business Document Understanding
by: Liu, Jiale, et al.
Published: (2024)
by: Liu, Jiale, et al.
Published: (2024)
Similar Items
-
Mango: Multi-Agent Web Navigation via Global-View Optimization
by: Tong, Weixi, et al.
Published: (2026) -
GraphicBench: A Planning Benchmark for Graphic Design with Language Agents
by: Ki, Dayeon, et al.
Published: (2025) -
Efficient Agent Training for Computer Use
by: He, Yanheng, et al.
Published: (2025) -
An Efficient Training Pipeline for Reasoning Graphical User Interface Agents
by: Pantazopoulos, Georgios, et al.
Published: (2025) -
CodeJudge: Evaluating Code Generation with Large Language Models
by: Tong, Weixi, et al.
Published: (2024)