Grounding Partially-Defined Events in Multimodal Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sanders, Kate, Kriz, Reno, Etter, David, Recknor, Hannah, Martin, Alexander, Carpenter, Cameron, Lin, Jingyang, Van Durme, Benjamin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912062219747328
author Sanders, Kate
Kriz, Reno
Etter, David
Recknor, Hannah
Martin, Alexander
Carpenter, Cameron
Lin, Jingyang
Van Durme, Benjamin
author_facet Sanders, Kate
Kriz, Reno
Etter, David
Recknor, Hannah
Martin, Alexander
Carpenter, Cameron
Lin, Jingyang
Van Durme, Benjamin
contents How are we able to learn about complex current events just from short snippets of video? While natural language enables straightforward ways to represent under-specified, partially observable events, visual data does not facilitate analogous methods and, consequently, introduces unique challenges in event understanding. With the growing prevalence of vision-capable AI agents, these systems must be able to model events from collections of unstructured video data. To tackle robust event modeling in multimodal settings, we introduce a multimodal formulation for partially-defined events and cast the extraction of these events as a three-stage span retrieval task. We propose a corresponding benchmark for this task, MultiVENT-G, that consists of 14.5 hours of densely annotated current event videos and 1,168 text documents, containing 22.8K labeled event-centric entities. We propose a collection of LLM-driven approaches to the task of multimodal event analysis, and evaluate them on MultiVENT-G. Results illustrate the challenges that abstract event understanding poses and demonstrates promise in event-centric video-language systems.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05267
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Grounding Partially-Defined Events in Multimodal Data
Sanders, Kate
Kriz, Reno
Etter, David
Recknor, Hannah
Martin, Alexander
Carpenter, Cameron
Lin, Jingyang
Van Durme, Benjamin
Computation and Language
Computer Vision and Pattern Recognition
How are we able to learn about complex current events just from short snippets of video? While natural language enables straightforward ways to represent under-specified, partially observable events, visual data does not facilitate analogous methods and, consequently, introduces unique challenges in event understanding. With the growing prevalence of vision-capable AI agents, these systems must be able to model events from collections of unstructured video data. To tackle robust event modeling in multimodal settings, we introduce a multimodal formulation for partially-defined events and cast the extraction of these events as a three-stage span retrieval task. We propose a corresponding benchmark for this task, MultiVENT-G, that consists of 14.5 hours of densely annotated current event videos and 1,168 text documents, containing 22.8K labeled event-centric entities. We propose a collection of LLM-driven approaches to the task of multimodal event analysis, and evaluate them on MultiVENT-G. Results illustrate the challenges that abstract event understanding poses and demonstrates promise in event-centric video-language systems.
title Grounding Partially-Defined Events in Multimodal Data
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.05267