VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Koh, Jing Yu, Lo, Robert, Jang, Lawrence, Duvvur, Vikram, Lim, Ming Chong, Huang, Po-Yu, Neubig, Graham, Zhou, Shuyan, Salakhutdinov, Ruslan, Fried, Daniel
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917685816721408
author Koh, Jing Yu
Lo, Robert
Jang, Lawrence
Duvvur, Vikram
Lim, Ming Chong
Huang, Po-Yu
Neubig, Graham
Zhou, Shuyan
Salakhutdinov, Ruslan
Fried, Daniel
author_facet Koh, Jing Yu
Lo, Robert
Jang, Lawrence
Duvvur, Vikram
Lim, Ming Chong
Huang, Po-Yu
Neubig, Graham
Zhou, Shuyan
Salakhutdinov, Ruslan
Fried, Daniel
contents Autonomous agents capable of planning, reasoning, and executing actions on the web offer a promising avenue for automating computer tasks. However, the majority of existing benchmarks primarily focus on text-based agents, neglecting many natural tasks that require visual information to effectively solve. Given that most computer interfaces cater to human perception, visual information often augments textual data in ways that text-only models struggle to harness effectively. To bridge this gap, we introduce VisualWebArena, a benchmark designed to assess the performance of multimodal web agents on realistic \textit{visually grounded tasks}. VisualWebArena comprises of a set of diverse and complex web-based tasks that evaluate various capabilities of autonomous multimodal agents. To perform on this benchmark, agents need to accurately process image-text inputs, interpret natural language instructions, and execute actions on websites to accomplish user-defined objectives. We conduct an extensive evaluation of state-of-the-art LLM-based autonomous agents, including several multimodal models. Through extensive quantitative and qualitative analysis, we identify several limitations of text-only LLM agents, and reveal gaps in the capabilities of state-of-the-art multimodal language agents. VisualWebArena provides a framework for evaluating multimodal autonomous language agents, and offers insights towards building stronger autonomous agents for the web. Our code, baseline models, and data is publicly available at https://jykoh.com/vwa.
format Preprint
id arxiv_https___arxiv_org_abs_2401_13649
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
Koh, Jing Yu
Lo, Robert
Jang, Lawrence
Duvvur, Vikram
Lim, Ming Chong
Huang, Po-Yu
Neubig, Graham
Zhou, Shuyan
Salakhutdinov, Ruslan
Fried, Daniel
Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
Autonomous agents capable of planning, reasoning, and executing actions on the web offer a promising avenue for automating computer tasks. However, the majority of existing benchmarks primarily focus on text-based agents, neglecting many natural tasks that require visual information to effectively solve. Given that most computer interfaces cater to human perception, visual information often augments textual data in ways that text-only models struggle to harness effectively. To bridge this gap, we introduce VisualWebArena, a benchmark designed to assess the performance of multimodal web agents on realistic \textit{visually grounded tasks}. VisualWebArena comprises of a set of diverse and complex web-based tasks that evaluate various capabilities of autonomous multimodal agents. To perform on this benchmark, agents need to accurately process image-text inputs, interpret natural language instructions, and execute actions on websites to accomplish user-defined objectives. We conduct an extensive evaluation of state-of-the-art LLM-based autonomous agents, including several multimodal models. Through extensive quantitative and qualitative analysis, we identify several limitations of text-only LLM agents, and reveal gaps in the capabilities of state-of-the-art multimodal language agents. VisualWebArena provides a framework for evaluating multimodal autonomous language agents, and offers insights towards building stronger autonomous agents for the web. Our code, baseline models, and data is publicly available at https://jykoh.com/vwa.
title VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
topic Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.13649