STEVE: A Step Verification Pipeline for Computer-use Agent Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Fanbin, Zhong, Zhisheng, Wei, Ziqin, Liu, Shu, Fu, Chi-Wing, Jia, Jiaya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913754787086336
author Lu, Fanbin
Zhong, Zhisheng
Wei, Ziqin
Liu, Shu
Fu, Chi-Wing
Jia, Jiaya
author_facet Lu, Fanbin
Zhong, Zhisheng
Wei, Ziqin
Liu, Shu
Fu, Chi-Wing
Jia, Jiaya
contents Developing AI agents to autonomously manipulate graphical user interfaces is a long challenging task. Recent advances in data scaling law inspire us to train computer-use agents with a scaled instruction set, yet using behavior cloning to train agents still requires immense high-quality trajectories. To meet the scalability need, we designed STEVE, a step verification pipeline for computer-use agent training. First, we establish a large instruction set for computer-use agents and collect trajectory data with some suboptimal agents. GPT-4o is used to verify the correctness of each step in the trajectories based on the screens before and after the action execution, assigning each step with a binary label. Last, we adopt the Kahneman and Tversky Optimization to optimize the agent from the binary stepwise labels. Extensive experiments manifest that our agent outperforms supervised finetuning by leveraging both positive and negative actions within a trajectory. Also, STEVE enables us to train a 7B vision-language model as a computer-use agent, achieving leading performance in the challenging live desktop environment WinAgentArena with great efficiency at a reduced cost. Code and data: https://github.com/FanbinLu/STEVE.
format Preprint
id arxiv_https___arxiv_org_abs_2503_12532
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle STEVE: A Step Verification Pipeline for Computer-use Agent Training
Lu, Fanbin
Zhong, Zhisheng
Wei, Ziqin
Liu, Shu
Fu, Chi-Wing
Jia, Jiaya
Computer Vision and Pattern Recognition
Artificial Intelligence
Developing AI agents to autonomously manipulate graphical user interfaces is a long challenging task. Recent advances in data scaling law inspire us to train computer-use agents with a scaled instruction set, yet using behavior cloning to train agents still requires immense high-quality trajectories. To meet the scalability need, we designed STEVE, a step verification pipeline for computer-use agent training. First, we establish a large instruction set for computer-use agents and collect trajectory data with some suboptimal agents. GPT-4o is used to verify the correctness of each step in the trajectories based on the screens before and after the action execution, assigning each step with a binary label. Last, we adopt the Kahneman and Tversky Optimization to optimize the agent from the binary stepwise labels. Extensive experiments manifest that our agent outperforms supervised finetuning by leveraging both positive and negative actions within a trajectory. Also, STEVE enables us to train a 7B vision-language model as a computer-use agent, achieving leading performance in the challenging live desktop environment WinAgentArena with great efficiency at a reduced cost. Code and data: https://github.com/FanbinLu/STEVE.
title STEVE: A Step Verification Pipeline for Computer-use Agent Training
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2503.12532