Autonomous Continual Learning for Environment Adaptation of Computer-Use Agents

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xue, Tianci, Liao, Zeyi, Shi, Tianneng, Wang, Zilu, Zhang, Kai, Song, Dawn, Su, Yu, Sun, Huan
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915999619481600
author Xue, Tianci
Liao, Zeyi
Shi, Tianneng
Wang, Zilu
Zhang, Kai
Song, Dawn
Su, Yu
Sun, Huan
author_facet Xue, Tianci
Liao, Zeyi
Shi, Tianneng
Wang, Zilu
Zhang, Kai
Song, Dawn
Su, Yu
Sun, Huan
contents Real-world digital environments are highly diverse and dynamic. These characteristics cause agents to frequently encounter unseen environments and distribution shifts, making continual learning in such environments essential for computer-use agents (CUAs). However, a key challenge lies in obtaining high-quality and environment-grounded training data without relying on costly human annotation. In this work, we introduce ACuRL, an Autonomous Curriculum Reinforcement Learning framework that continually adapts agents to specific environments with zero human data. The agent first explores an environment to acquire initial experiences. During subsequent iterative training, a curriculum task generator leverages these experiences together with feedback from the previous iteration to synthesize new tasks tailored for the agent's current capabilities. To provide reliable reward signals, we introduce CUAJudge, a robust automatic evaluator for CUAs that achieves 93% agreement with human judgments. Empirically, our method effectively enables both intra-environment and cross-environment continual learning, yielding 3-29% absolute performance gains on the target environments without catastrophic forgetting on others. We also show that it can mitigate performance degradation under environment changes (e.g., version updates, platform migration, and resolution shifts). Further analyses show highly sparse updates (e.g., only 20% parameters), which helps explain the effective and robust adaptation.
format Preprint
id arxiv_https___arxiv_org_abs_2602_10356
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Autonomous Continual Learning for Environment Adaptation of Computer-Use Agents
Xue, Tianci
Liao, Zeyi
Shi, Tianneng
Wang, Zilu
Zhang, Kai
Song, Dawn
Su, Yu
Sun, Huan
Computation and Language
Real-world digital environments are highly diverse and dynamic. These characteristics cause agents to frequently encounter unseen environments and distribution shifts, making continual learning in such environments essential for computer-use agents (CUAs). However, a key challenge lies in obtaining high-quality and environment-grounded training data without relying on costly human annotation. In this work, we introduce ACuRL, an Autonomous Curriculum Reinforcement Learning framework that continually adapts agents to specific environments with zero human data. The agent first explores an environment to acquire initial experiences. During subsequent iterative training, a curriculum task generator leverages these experiences together with feedback from the previous iteration to synthesize new tasks tailored for the agent's current capabilities. To provide reliable reward signals, we introduce CUAJudge, a robust automatic evaluator for CUAs that achieves 93% agreement with human judgments. Empirically, our method effectively enables both intra-environment and cross-environment continual learning, yielding 3-29% absolute performance gains on the target environments without catastrophic forgetting on others. We also show that it can mitigate performance degradation under environment changes (e.g., version updates, platform migration, and resolution shifts). Further analyses show highly sparse updates (e.g., only 20% parameters), which helps explain the effective and robust adaptation.
title Autonomous Continual Learning for Environment Adaptation of Computer-Use Agents
topic Computation and Language
url https://arxiv.org/abs/2602.10356