ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform Data
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908546381119488 |
|---|---|
| author | Liu, Zhaoyang Xie, Jingjing Ding, Zichen Li, Zehao Yang, Bowen Wu, Zhenyu Wang, Xuehui Sun, Qiushi Liu, Shi Wang, Weiyun Ye, Shenglong Li, Qingyun Dong, Xuan Yu, Yue Lu, Chenyu Mo, YunXiang Yan, Yao Tian, Zeyue Zhang, Xiao Huang, Yuan Liu, Yiqian Su, Weijie Luo, Gen Yue, Xiangyu Qi, Biqing Chen, Kai Zhou, Bowen Qiao, Yu Chen, Qifeng Wang, Wenhai |
| author_facet | Liu, Zhaoyang Xie, Jingjing Ding, Zichen Li, Zehao Yang, Bowen Wu, Zhenyu Wang, Xuehui Sun, Qiushi Liu, Shi Wang, Weiyun Ye, Shenglong Li, Qingyun Dong, Xuan Yu, Yue Lu, Chenyu Mo, YunXiang Yan, Yao Tian, Zeyue Zhang, Xiao Huang, Yuan Liu, Yiqian Su, Weijie Luo, Gen Yue, Xiangyu Qi, Biqing Chen, Kai Zhou, Bowen Qiao, Yu Chen, Qifeng Wang, Wenhai |
| contents | Vision-Language Models (VLMs) have enabled computer use agents (CUAs) that operate GUIs autonomously, showing great potential, yet progress is limited by the lack of large-scale, open-source computer use data and foundation models. In this work, we introduce ScaleCUA, a step toward scaling open-source CUAs. It offers a large-scale dataset spanning 6 operating systems and 3 task domains, built via a closed-loop pipeline uniting automated agents with human experts. Trained on this scaled-up data, ScaleCUA can operate seamlessly across platforms. Specifically, it delivers strong gains over baselines (+26.6 on WebArena-Lite-v2, +10.7 on ScreenSpot-Pro) and sets new state-of-the-art results (94.4% on MMBench-GUI L1-Hard, 60.6% on OSWorld-G, 47.4% on WebArena-Lite-v2). These findings underscore the power of data-driven scaling for general-purpose computer use agents. We will release data, models, and code to advance future research: https://github.com/OpenGVLab/ScaleCUA. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_15221 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform Data Liu, Zhaoyang Xie, Jingjing Ding, Zichen Li, Zehao Yang, Bowen Wu, Zhenyu Wang, Xuehui Sun, Qiushi Liu, Shi Wang, Weiyun Ye, Shenglong Li, Qingyun Dong, Xuan Yu, Yue Lu, Chenyu Mo, YunXiang Yan, Yao Tian, Zeyue Zhang, Xiao Huang, Yuan Liu, Yiqian Su, Weijie Luo, Gen Yue, Xiangyu Qi, Biqing Chen, Kai Zhou, Bowen Qiao, Yu Chen, Qifeng Wang, Wenhai Computer Vision and Pattern Recognition Vision-Language Models (VLMs) have enabled computer use agents (CUAs) that operate GUIs autonomously, showing great potential, yet progress is limited by the lack of large-scale, open-source computer use data and foundation models. In this work, we introduce ScaleCUA, a step toward scaling open-source CUAs. It offers a large-scale dataset spanning 6 operating systems and 3 task domains, built via a closed-loop pipeline uniting automated agents with human experts. Trained on this scaled-up data, ScaleCUA can operate seamlessly across platforms. Specifically, it delivers strong gains over baselines (+26.6 on WebArena-Lite-v2, +10.7 on ScreenSpot-Pro) and sets new state-of-the-art results (94.4% on MMBench-GUI L1-Hard, 60.6% on OSWorld-G, 47.4% on WebArena-Lite-v2). These findings underscore the power of data-driven scaling for general-purpose computer use agents. We will release data, models, and code to advance future research: https://github.com/OpenGVLab/ScaleCUA. |
| title | ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform Data |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2509.15221 |