WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Qi, Zehan, Liu, Xiao, Iong, Iat Long, Lai, Hanyu, Sun, Xueqiao, Zhao, Wenyi, Yang, Yu, Yang, Xinyue, Sun, Jiadai, Yao, Shuntian, Zhang, Tianjie, Xu, Wei, Tang, Jie, Dong, Yuxiao
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910800726196224
author Qi, Zehan
Liu, Xiao
Iong, Iat Long
Lai, Hanyu
Sun, Xueqiao
Zhao, Wenyi
Yang, Yu
Yang, Xinyue
Sun, Jiadai
Yao, Shuntian
Zhang, Tianjie
Xu, Wei
Tang, Jie
Dong, Yuxiao
author_facet Qi, Zehan
Liu, Xiao
Iong, Iat Long
Lai, Hanyu
Sun, Xueqiao
Zhao, Wenyi
Yang, Yu
Yang, Xinyue
Sun, Jiadai
Yao, Shuntian
Zhang, Tianjie
Xu, Wei
Tang, Jie
Dong, Yuxiao
contents Large language models (LLMs) have shown remarkable potential as autonomous agents, particularly in web-based tasks. However, existing LLM web agents heavily rely on expensive proprietary LLM APIs, while open LLMs lack the necessary decision-making capabilities. This paper introduces WebRL, a self-evolving online curriculum reinforcement learning framework designed to train high-performance web agents using open LLMs. WebRL addresses three key challenges in building LLM web agents, including the scarcity of training tasks, sparse feedback signals, and policy distribution drift in online learning. Specifically, WebRL incorporates 1) a self-evolving curriculum that generates new tasks from unsuccessful attempts, 2) a robust outcome-supervised reward model (ORM), and 3) adaptive reinforcement learning strategies to ensure consistent improvements. We apply WebRL to transform open Llama-3.1 and GLM-4 models into proficient web agents. On WebArena-Lite, WebRL improves the success rate of Llama-3.1-8B from 4.8% to 42.4%, and from 6.1% to 43% for GLM-4-9B. These open models significantly surpass the performance of GPT-4-Turbo (17.6%) and GPT-4o (13.9%) and outperform previous state-of-the-art web agents trained on open LLMs (AutoWebGLM, 18.2%). Our findings demonstrate WebRL's effectiveness in bridging the gap between open and proprietary LLM-based web agents, paving the way for more accessible and powerful autonomous web interaction systems.
format Preprint
id arxiv_https___arxiv_org_abs_2411_02337
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning
Qi, Zehan
Liu, Xiao
Iong, Iat Long
Lai, Hanyu
Sun, Xueqiao
Zhao, Wenyi
Yang, Yu
Yang, Xinyue
Sun, Jiadai
Yao, Shuntian
Zhang, Tianjie
Xu, Wei
Tang, Jie
Dong, Yuxiao
Computation and Language
Large language models (LLMs) have shown remarkable potential as autonomous agents, particularly in web-based tasks. However, existing LLM web agents heavily rely on expensive proprietary LLM APIs, while open LLMs lack the necessary decision-making capabilities. This paper introduces WebRL, a self-evolving online curriculum reinforcement learning framework designed to train high-performance web agents using open LLMs. WebRL addresses three key challenges in building LLM web agents, including the scarcity of training tasks, sparse feedback signals, and policy distribution drift in online learning. Specifically, WebRL incorporates 1) a self-evolving curriculum that generates new tasks from unsuccessful attempts, 2) a robust outcome-supervised reward model (ORM), and 3) adaptive reinforcement learning strategies to ensure consistent improvements. We apply WebRL to transform open Llama-3.1 and GLM-4 models into proficient web agents. On WebArena-Lite, WebRL improves the success rate of Llama-3.1-8B from 4.8% to 42.4%, and from 6.1% to 43% for GLM-4-9B. These open models significantly surpass the performance of GPT-4-Turbo (17.6%) and GPT-4o (13.9%) and outperform previous state-of-the-art web agents trained on open LLMs (AutoWebGLM, 18.2%). Our findings demonstrate WebRL's effectiveness in bridging the gap between open and proprietary LLM-based web agents, paving the way for more accessible and powerful autonomous web interaction systems.
title WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning
topic Computation and Language
url https://arxiv.org/abs/2411.02337