Covering Human Action Space for Computer Use: Data Synthesis and Benchmark

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Miaosen, Zhao, Xiaohan, Tan, Zhihong, Huoshen, Zhou, Fan, Yijia, Yang, Yifan, Qiu, Kai, Liu, Bei, Wagle, Justin, Yin, Chenzhong, Cheng, Mingxi, Li, Ji, Dai, Qi, Luo, Chong, Yang, Xu, Geng, Xin, Guo, Baining
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914559864864768
author Zhang, Miaosen
Zhao, Xiaohan
Tan, Zhihong
Huoshen, Zhou
Fan, Yijia
Yang, Yifan
Qiu, Kai
Liu, Bei
Wagle, Justin
Yin, Chenzhong
Cheng, Mingxi
Li, Ji
Dai, Qi
Luo, Chong
Yang, Xu
Geng, Xin
Guo, Baining
author_facet Zhang, Miaosen
Zhao, Xiaohan
Tan, Zhihong
Huoshen, Zhou
Fan, Yijia
Yang, Yifan
Qiu, Kai
Liu, Bei
Wagle, Justin
Yin, Chenzhong
Cheng, Mingxi
Li, Ji
Dai, Qi
Luo, Chong
Yang, Xu
Geng, Xin
Guo, Baining
contents Computer-use agents (CUAs) automate on-screen work, as illustrated by GPT-5.4 and Claude. Yet their reliability on complex, low-frequency interactions is still poor, limiting user trust. Our analysis of failure cases from advanced models suggests a long-tail pattern in GUI operations, where a relatively small fraction of complex and diverse interactions accounts for a disproportionate share of task failures. We hypothesize that this issue largely stems from the scarcity of data for complex interactions. To address this problem, we propose a new benchmark CUActSpot for evaluating models' capabilities on complex interactions across five modalities: GUI, text, table, canvas, and natural image, as well as a variety of actions (click, drag, draw, etc.), covering a broader range of interaction types than prior click-centric benchmarks that focus mainly on GUI widgets. We also design a renderer-based data-synthesis pipeline: scenes are automatically generated for each modality, screenshots and element coordinates are recorded, and an LLM produces matching instructions and action traces. After training on this corpus, our Phi-Ground-Any-4B outperforms open-source models with fewer than 32B parameters. We will release our benchmark, data, code, and models at https://github.com/microsoft/Phi-Ground.git
format Preprint
id arxiv_https___arxiv_org_abs_2605_12501
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Covering Human Action Space for Computer Use: Data Synthesis and Benchmark
Zhang, Miaosen
Zhao, Xiaohan
Tan, Zhihong
Huoshen, Zhou
Fan, Yijia
Yang, Yifan
Qiu, Kai
Liu, Bei
Wagle, Justin
Yin, Chenzhong
Cheng, Mingxi
Li, Ji
Dai, Qi
Luo, Chong
Yang, Xu
Geng, Xin
Guo, Baining
Computer Vision and Pattern Recognition
Computer-use agents (CUAs) automate on-screen work, as illustrated by GPT-5.4 and Claude. Yet their reliability on complex, low-frequency interactions is still poor, limiting user trust. Our analysis of failure cases from advanced models suggests a long-tail pattern in GUI operations, where a relatively small fraction of complex and diverse interactions accounts for a disproportionate share of task failures. We hypothesize that this issue largely stems from the scarcity of data for complex interactions. To address this problem, we propose a new benchmark CUActSpot for evaluating models' capabilities on complex interactions across five modalities: GUI, text, table, canvas, and natural image, as well as a variety of actions (click, drag, draw, etc.), covering a broader range of interaction types than prior click-centric benchmarks that focus mainly on GUI widgets. We also design a renderer-based data-synthesis pipeline: scenes are automatically generated for each modality, screenshots and element coordinates are recorded, and an LLM produces matching instructions and action traces. After training on this corpus, our Phi-Ground-Any-4B outperforms open-source models with fewer than 32B parameters. We will release our benchmark, data, code, and models at https://github.com/microsoft/Phi-Ground.git
title Covering Human Action Space for Computer Use: Data Synthesis and Benchmark
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.12501