Large Language Models Can Self-Improve At Web Agent Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Patel, Ajay, Hofmarcher, Markus, Leoveanu-Condrei, Claudiu, Dinu, Marius-Constantin, Callison-Burch, Chris, Hochreiter, Sepp
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909332058144768
author Patel, Ajay
Hofmarcher, Markus
Leoveanu-Condrei, Claudiu
Dinu, Marius-Constantin
Callison-Burch, Chris
Hochreiter, Sepp
author_facet Patel, Ajay
Hofmarcher, Markus
Leoveanu-Condrei, Claudiu
Dinu, Marius-Constantin
Callison-Burch, Chris
Hochreiter, Sepp
contents Training models to act as agents that can effectively navigate and perform actions in a complex environment, such as a web browser, has typically been challenging due to lack of training data. Large language models (LLMs) have recently demonstrated some capability to navigate novel environments as agents in a zero-shot or few-shot fashion, purely guided by natural language instructions as prompts. Recent research has also demonstrated LLMs have the capability to exceed their base performance through self-improvement, i.e. fine-tuning on data generated by the model itself. In this work, we explore the extent to which LLMs can self-improve their performance as agents in long-horizon tasks in a complex environment using the WebArena benchmark. In WebArena, an agent must autonomously navigate and perform actions on web pages to achieve a specified objective. We explore fine-tuning on three distinct synthetic training data mixtures and achieve a 31\% improvement in task completion rate over the base model on the WebArena benchmark through a self-improvement procedure. We additionally contribute novel evaluation metrics for assessing the performance, robustness, capabilities, and quality of trajectories of our fine-tuned agent models to a greater degree than simple, aggregate-level benchmark scores currently used to measure self-improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20309
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Large Language Models Can Self-Improve At Web Agent Tasks
Patel, Ajay
Hofmarcher, Markus
Leoveanu-Condrei, Claudiu
Dinu, Marius-Constantin
Callison-Burch, Chris
Hochreiter, Sepp
Machine Learning
Artificial Intelligence
Computation and Language
Training models to act as agents that can effectively navigate and perform actions in a complex environment, such as a web browser, has typically been challenging due to lack of training data. Large language models (LLMs) have recently demonstrated some capability to navigate novel environments as agents in a zero-shot or few-shot fashion, purely guided by natural language instructions as prompts. Recent research has also demonstrated LLMs have the capability to exceed their base performance through self-improvement, i.e. fine-tuning on data generated by the model itself. In this work, we explore the extent to which LLMs can self-improve their performance as agents in long-horizon tasks in a complex environment using the WebArena benchmark. In WebArena, an agent must autonomously navigate and perform actions on web pages to achieve a specified objective. We explore fine-tuning on three distinct synthetic training data mixtures and achieve a 31\% improvement in task completion rate over the base model on the WebArena benchmark through a self-improvement procedure. We additionally contribute novel evaluation metrics for assessing the performance, robustness, capabilities, and quality of trajectories of our fine-tuned agent models to a greater degree than simple, aggregate-level benchmark scores currently used to measure self-improvement.
title Large Language Models Can Self-Improve At Web Agent Tasks
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2405.20309