Efficient Agent Training for Computer Use

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: He, Yanheng, Jin, Jiahe, Liu, Pengfei
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911480406867968
author He, Yanheng
Jin, Jiahe
Liu, Pengfei
author_facet He, Yanheng
Jin, Jiahe
Liu, Pengfei
contents Scaling up high-quality trajectory data has long been a critical bottleneck for developing human-like computer use agents. We introduce PC Agent-E, an efficient agent training framework that significantly reduces reliance on large-scale human demonstrations. Starting with just 312 human-annotated computer use trajectories, we further augment them by synthesizing diverse alternative action decisions with Claude 3.7 Sonnet. Trained on these enriched trajectories, our PC Agent-E model achieved a remarkable 141 relative improvement, and even surpassed the Claude 3.7 Sonnet by 10% in relative terms on WindowsAgentArena-V2, an improved benchmark we also released. By integrating robust human computer use skills with automated AI data synthesis capabilities, our method not only brought substantial improvements over training on human trajectories alone, but also significantly surpassed direct distillation from Claude 3.7 Sonnet. Code, data and models are available at https://github.com/GAIR-NLP/PC-Agent-E
format Preprint
id arxiv_https___arxiv_org_abs_2505_13909
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Agent Training for Computer Use
He, Yanheng
Jin, Jiahe
Liu, Pengfei
Artificial Intelligence
Computation and Language
Machine Learning
Scaling up high-quality trajectory data has long been a critical bottleneck for developing human-like computer use agents. We introduce PC Agent-E, an efficient agent training framework that significantly reduces reliance on large-scale human demonstrations. Starting with just 312 human-annotated computer use trajectories, we further augment them by synthesizing diverse alternative action decisions with Claude 3.7 Sonnet. Trained on these enriched trajectories, our PC Agent-E model achieved a remarkable 141 relative improvement, and even surpassed the Claude 3.7 Sonnet by 10% in relative terms on WindowsAgentArena-V2, an improved benchmark we also released. By integrating robust human computer use skills with automated AI data synthesis capabilities, our method not only brought substantial improvements over training on human trajectories alone, but also significantly surpassed direct distillation from Claude 3.7 Sonnet. Code, data and models are available at https://github.com/GAIR-NLP/PC-Agent-E
title Efficient Agent Training for Computer Use
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.13909