UI-Evol: Automatic Knowledge Evolving for Computer Use Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866911245209174016 |
|---|---|
| author | Zhang, Ziyun Liu, Xinyi Zhang, Xiaoyi Wang, Jun Chen, Gang Lu, Yan |
| author_facet | Zhang, Ziyun Liu, Xinyi Zhang, Xiaoyi Wang, Jun Chen, Gang Lu, Yan |
| contents | External knowledge has played a crucial role in the recent development of computer use agents. We identify a critical knowledge-execution gap: retrieved knowledge often fails to translate into effective real-world task execution. Our analysis shows even 90% correct knowledge yields only 41% execution success rate. To bridge this gap, we propose UI-Evol, a plug-and-play module for autonomous GUI knowledge evolution. UI-Evol consists of two stages: a Retrace Stage that extracts faithful objective action sequences from actual agent-environment interactions, and a Critique Stage that refines existing knowledge by comparing these sequences against external references. We conduct comprehensive experiments on the OSWorld benchmark with the state-of-the-art Agent S2. Our results demonstrate that UI-Evol not only significantly boosts task performance but also addresses a previously overlooked issue of high behavioral standard deviation in computer use agents, leading to superior performance on computer use tasks and substantially improved agent reliability. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_21964 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | UI-Evol: Automatic Knowledge Evolving for Computer Use Agents Zhang, Ziyun Liu, Xinyi Zhang, Xiaoyi Wang, Jun Chen, Gang Lu, Yan Human-Computer Interaction Computation and Language External knowledge has played a crucial role in the recent development of computer use agents. We identify a critical knowledge-execution gap: retrieved knowledge often fails to translate into effective real-world task execution. Our analysis shows even 90% correct knowledge yields only 41% execution success rate. To bridge this gap, we propose UI-Evol, a plug-and-play module for autonomous GUI knowledge evolution. UI-Evol consists of two stages: a Retrace Stage that extracts faithful objective action sequences from actual agent-environment interactions, and a Critique Stage that refines existing knowledge by comparing these sequences against external references. We conduct comprehensive experiments on the OSWorld benchmark with the state-of-the-art Agent S2. Our results demonstrate that UI-Evol not only significantly boosts task performance but also addresses a previously overlooked issue of high behavioral standard deviation in computer use agents, leading to superior performance on computer use tasks and substantially improved agent reliability. |
| title | UI-Evol: Automatic Knowledge Evolving for Computer Use Agents |
| topic | Human-Computer Interaction Computation and Language |
| url | https://arxiv.org/abs/2505.21964 |