The BrowserGym Ecosystem for Web Agent Research

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: De Chezelles, Thibault Le Sellier, Gasse, Maxime, Drouin, Alexandre, Caccia, Massimo, Boisvert, Léo, Thakkar, Megh, Marty, Tom, Assouel, Rim, Shayegan, Sahar Omidi, Jang, Lawrence Keunho, Lù, Xing Han, Yoran, Ori, Kong, Dehan, Xu, Frank F., Reddy, Siva, Cappart, Quentin, Neubig, Graham, Salakhutdinov, Ruslan, Chapados, Nicolas, Lacoste, Alexandre
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909516045484032
author De Chezelles, Thibault Le Sellier
Gasse, Maxime
Drouin, Alexandre
Caccia, Massimo
Boisvert, Léo
Thakkar, Megh
Marty, Tom
Assouel, Rim
Shayegan, Sahar Omidi
Jang, Lawrence Keunho
Lù, Xing Han
Yoran, Ori
Kong, Dehan
Xu, Frank F.
Reddy, Siva
Cappart, Quentin
Neubig, Graham
Salakhutdinov, Ruslan
Chapados, Nicolas
Lacoste, Alexandre
author_facet De Chezelles, Thibault Le Sellier
Gasse, Maxime
Drouin, Alexandre
Caccia, Massimo
Boisvert, Léo
Thakkar, Megh
Marty, Tom
Assouel, Rim
Shayegan, Sahar Omidi
Jang, Lawrence Keunho
Lù, Xing Han
Yoran, Ori
Kong, Dehan
Xu, Frank F.
Reddy, Siva
Cappart, Quentin
Neubig, Graham
Salakhutdinov, Ruslan
Chapados, Nicolas
Lacoste, Alexandre
contents The BrowserGym ecosystem addresses the growing need for efficient evaluation and benchmarking of web agents, particularly those leveraging automation and Large Language Models (LLMs). Many existing benchmarks suffer from fragmentation and inconsistent evaluation methodologies, making it challenging to achieve reliable comparisons and reproducible results. In an earlier work, Drouin et al. (2024) introduced BrowserGym which aims to solve this by providing a unified, gym-like environment with well-defined observation and action spaces, facilitating standardized evaluation across diverse benchmarks. We propose an extended BrowserGym-based ecosystem for web agent research, which unifies existing benchmarks from the literature and includes AgentLab, a complementary framework that aids in agent creation, testing, and analysis. Our proposed ecosystem offers flexibility for integrating new benchmarks while ensuring consistent evaluation and comprehensive experiment management. As a supporting evidence, we conduct the first large-scale, multi-benchmark web agent experiment and compare the performance of 6 state-of-the-art LLMs across 6 popular web agent benchmarks made available in BrowserGym. Among other findings, our results highlight a large discrepancy between OpenAI and Anthropic's latests models, with Claude-3.5-Sonnet leading the way on almost all benchmarks, except on vision-related tasks where GPT-4o is superior. Despite these advancements, our results emphasize that building robust and efficient web agents remains a significant challenge, due to the inherent complexity of real-world web environments and the limitations of current models.
format Preprint
id arxiv_https___arxiv_org_abs_2412_05467
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The BrowserGym Ecosystem for Web Agent Research
De Chezelles, Thibault Le Sellier
Gasse, Maxime
Drouin, Alexandre
Caccia, Massimo
Boisvert, Léo
Thakkar, Megh
Marty, Tom
Assouel, Rim
Shayegan, Sahar Omidi
Jang, Lawrence Keunho
Lù, Xing Han
Yoran, Ori
Kong, Dehan
Xu, Frank F.
Reddy, Siva
Cappart, Quentin
Neubig, Graham
Salakhutdinov, Ruslan
Chapados, Nicolas
Lacoste, Alexandre
Machine Learning
Artificial Intelligence
Software Engineering
The BrowserGym ecosystem addresses the growing need for efficient evaluation and benchmarking of web agents, particularly those leveraging automation and Large Language Models (LLMs). Many existing benchmarks suffer from fragmentation and inconsistent evaluation methodologies, making it challenging to achieve reliable comparisons and reproducible results. In an earlier work, Drouin et al. (2024) introduced BrowserGym which aims to solve this by providing a unified, gym-like environment with well-defined observation and action spaces, facilitating standardized evaluation across diverse benchmarks. We propose an extended BrowserGym-based ecosystem for web agent research, which unifies existing benchmarks from the literature and includes AgentLab, a complementary framework that aids in agent creation, testing, and analysis. Our proposed ecosystem offers flexibility for integrating new benchmarks while ensuring consistent evaluation and comprehensive experiment management. As a supporting evidence, we conduct the first large-scale, multi-benchmark web agent experiment and compare the performance of 6 state-of-the-art LLMs across 6 popular web agent benchmarks made available in BrowserGym. Among other findings, our results highlight a large discrepancy between OpenAI and Anthropic's latests models, with Claude-3.5-Sonnet leading the way on almost all benchmarks, except on vision-related tasks where GPT-4o is superior. Despite these advancements, our results emphasize that building robust and efficient web agents remains a significant challenge, due to the inherent complexity of real-world web environments and the limitations of current models.
title The BrowserGym Ecosystem for Web Agent Research
topic Machine Learning
Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2412.05467