VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jang, Lawrence, Li, Yinheng, Zhao, Dan, Ding, Charles, Lin, Justin, Liang, Paul Pu, Bonatti, Rogerio, Koishida, Kazuhito
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916616361476096
author Jang, Lawrence
Li, Yinheng
Zhao, Dan
Ding, Charles
Lin, Justin
Liang, Paul Pu
Bonatti, Rogerio
Koishida, Kazuhito
author_facet Jang, Lawrence
Li, Yinheng
Zhao, Dan
Ding, Charles
Lin, Justin
Liang, Paul Pu
Bonatti, Rogerio
Koishida, Kazuhito
contents Videos are often used to learn or extract the necessary information to complete tasks in ways different than what text and static imagery alone can provide. However, many existing agent benchmarks neglect long-context video understanding, instead focusing on text or static image inputs. To bridge this gap, we introduce VideoWebArena (VideoWA), a benchmark for evaluating the capabilities of long-context multimodal agents for video understanding. VideoWA consists of 2,021 web agent tasks based on manually crafted video tutorials, which total almost four hours of content. For our benchmark, we define a taxonomy of long-context video-based agent tasks with two main areas of focus: skill retention and factual retention. While skill retention tasks evaluate whether an agent can use a given human demonstration to complete a task efficiently, the factual retention task evaluates whether an agent can retrieve instruction-relevant information from a video to complete a task. We find that the best model achieves 13.3% success on factual retention tasks and 45.8% on factual retention QA pairs, far below human performance at 73.9% and 79.3%, respectively. On skill retention tasks, long-context models perform worse with tutorials than without, exhibiting a 5% performance decrease in WebArena tasks and a 10.3% decrease in VisualWebArena tasks. Our work highlights the need to improve the agentic abilities of long-context multimodal models and provides a testbed for future development with long-context video agents.
format Preprint
id arxiv_https___arxiv_org_abs_2410_19100
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks
Jang, Lawrence
Li, Yinheng
Zhao, Dan
Ding, Charles
Lin, Justin
Liang, Paul Pu
Bonatti, Rogerio
Koishida, Kazuhito
Computer Vision and Pattern Recognition
Artificial Intelligence
Videos are often used to learn or extract the necessary information to complete tasks in ways different than what text and static imagery alone can provide. However, many existing agent benchmarks neglect long-context video understanding, instead focusing on text or static image inputs. To bridge this gap, we introduce VideoWebArena (VideoWA), a benchmark for evaluating the capabilities of long-context multimodal agents for video understanding. VideoWA consists of 2,021 web agent tasks based on manually crafted video tutorials, which total almost four hours of content. For our benchmark, we define a taxonomy of long-context video-based agent tasks with two main areas of focus: skill retention and factual retention. While skill retention tasks evaluate whether an agent can use a given human demonstration to complete a task efficiently, the factual retention task evaluates whether an agent can retrieve instruction-relevant information from a video to complete a task. We find that the best model achieves 13.3% success on factual retention tasks and 45.8% on factual retention QA pairs, far below human performance at 73.9% and 79.3%, respectively. On skill retention tasks, long-context models perform worse with tutorials than without, exhibiting a 5% performance decrease in WebArena tasks and a 10.3% decrease in VisualWebArena tasks. Our work highlights the need to improve the agentic abilities of long-context multimodal models and provides a testbed for future development with long-context video agents.
title VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2410.19100