WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wei, Zhepei, Yao, Wenlin, Liu, Yao, Zhang, Weizhi, Lu, Qin, Qiu, Liang, Yu, Changlong, Xu, Puyang, Zhang, Chao, Yin, Bing, Yun, Hyokun, Li, Lihong
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909831955218432
author Wei, Zhepei
Yao, Wenlin
Liu, Yao
Zhang, Weizhi
Lu, Qin
Qiu, Liang
Yu, Changlong
Xu, Puyang
Zhang, Chao
Yin, Bing
Yun, Hyokun
Li, Lihong
author_facet Wei, Zhepei
Yao, Wenlin
Liu, Yao
Zhang, Weizhi
Lu, Qin
Qiu, Liang
Yu, Changlong
Xu, Puyang
Zhang, Chao
Yin, Bing
Yun, Hyokun
Li, Lihong
contents While reinforcement learning (RL) has demonstrated remarkable success in enhancing large language models (LLMs), it has primarily focused on single-turn tasks such as solving math problems. Training effective web agents for multi-turn interactions remains challenging due to the complexity of long-horizon decision-making across dynamic web interfaces. In this work, we present WebAgent-R1, a simple yet effective end-to-end multi-turn RL framework for training web agents. It learns directly from online interactions with web environments by asynchronously generating diverse trajectories, entirely guided by binary rewards depending on task success. Experiments on the WebArena-Lite benchmark demonstrate the effectiveness of WebAgent-R1, boosting the task success rate of Qwen-2.5-3B from 6.1% to 33.9% and Llama-3.1-8B from 8.5% to 44.8%, significantly outperforming existing state-of-the-art methods and strong proprietary models such as OpenAI o3. In-depth analyses reveal the effectiveness of the thinking-based prompting strategy and test-time scaling through increased interactions for web tasks. We further investigate different RL initialization policies by introducing two variants, namely WebAgent-R1-Zero and WebAgent-R1-CoT, which highlight the importance of the warm-up training stage (i.e., behavior cloning) and provide insights on incorporating long chain-of-thought (CoT) reasoning in web agents.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16421
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning
Wei, Zhepei
Yao, Wenlin
Liu, Yao
Zhang, Weizhi
Lu, Qin
Qiu, Liang
Yu, Changlong
Xu, Puyang
Zhang, Chao
Yin, Bing
Yun, Hyokun
Li, Lihong
Computation and Language
Machine Learning
While reinforcement learning (RL) has demonstrated remarkable success in enhancing large language models (LLMs), it has primarily focused on single-turn tasks such as solving math problems. Training effective web agents for multi-turn interactions remains challenging due to the complexity of long-horizon decision-making across dynamic web interfaces. In this work, we present WebAgent-R1, a simple yet effective end-to-end multi-turn RL framework for training web agents. It learns directly from online interactions with web environments by asynchronously generating diverse trajectories, entirely guided by binary rewards depending on task success. Experiments on the WebArena-Lite benchmark demonstrate the effectiveness of WebAgent-R1, boosting the task success rate of Qwen-2.5-3B from 6.1% to 33.9% and Llama-3.1-8B from 8.5% to 44.8%, significantly outperforming existing state-of-the-art methods and strong proprietary models such as OpenAI o3. In-depth analyses reveal the effectiveness of the thinking-based prompting strategy and test-time scaling through increased interactions for web tasks. We further investigate different RL initialization policies by introducing two variants, namely WebAgent-R1-Zero and WebAgent-R1-CoT, which highlight the importance of the warm-up training stage (i.e., behavior cloning) and provide insights on incorporating long chain-of-thought (CoT) reasoning in web agents.
title WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.16421