On the Reliability of Computer Use Agents

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gonzalez-Pumariega, Gonzalo, Agashe, Saaket, Yang, Jiachen, Li, Ang, Wang, Xin Eric
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908978663915520
author Gonzalez-Pumariega, Gonzalo
Agashe, Saaket
Yang, Jiachen
Li, Ang
Wang, Xin Eric
author_facet Gonzalez-Pumariega, Gonzalo
Agashe, Saaket
Yang, Jiachen
Li, Ang
Wang, Xin Eric
contents Computer-use agents have rapidly improved on real-world tasks such as web navigation, desktop automation, and software interaction, in some cases surpassing human performance. Yet even when the task and model are unchanged, an agent that succeeds once may fail on a repeated execution of the same task. This raises a fundamental question: if an agent can succeed at a task once, what prevents it from doing so reliably? In this work, we study the sources of unreliability in computer-use agents through three factors: stochasticity during execution, ambiguity in task specification, and variability in agent behavior. We analyze these factors on OSWorld using repeated executions of the same task together with paired statistical tests that capture task-level changes across settings. Our analysis shows that reliability depends on both how tasks are specified and how agent behavior varies across executions. These findings suggest the need to evaluate agents under repeated execution, to allow agents to resolve task ambiguity through interaction, and to favor strategies that remain stable across runs.
format Preprint
id arxiv_https___arxiv_org_abs_2604_17849
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle On the Reliability of Computer Use Agents
Gonzalez-Pumariega, Gonzalo
Agashe, Saaket
Yang, Jiachen
Li, Ang
Wang, Xin Eric
Artificial Intelligence
Computer-use agents have rapidly improved on real-world tasks such as web navigation, desktop automation, and software interaction, in some cases surpassing human performance. Yet even when the task and model are unchanged, an agent that succeeds once may fail on a repeated execution of the same task. This raises a fundamental question: if an agent can succeed at a task once, what prevents it from doing so reliably? In this work, we study the sources of unreliability in computer-use agents through three factors: stochasticity during execution, ambiguity in task specification, and variability in agent behavior. We analyze these factors on OSWorld using repeated executions of the same task together with paired statistical tests that capture task-level changes across settings. Our analysis shows that reliability depends on both how tasks are specified and how agent behavior varies across executions. These findings suggest the need to evaluate agents under repeated execution, to allow agents to resolve task ambiguity through interaction, and to favor strategies that remain stable across runs.
title On the Reliability of Computer Use Agents
topic Artificial Intelligence
url https://arxiv.org/abs/2604.17849