JoVALE: Detecting Human Actions in Video Using Audiovisual and Language Contexts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Son, Taein, Seo, Soo Won, Kim, Jisong, Lee, Seok Hwan, Choi, Jun Won
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912216325816320
author Son, Taein
Seo, Soo Won
Kim, Jisong
Lee, Seok Hwan
Choi, Jun Won
author_facet Son, Taein
Seo, Soo Won
Kim, Jisong
Lee, Seok Hwan
Choi, Jun Won
contents Video Action Detection (VAD) entails localizing and categorizing action instances within videos, which inherently consist of diverse information sources such as audio, visual cues, and surrounding scene contexts. Leveraging this multi-modal information effectively for VAD poses a significant challenge, as the model must identify action-relevant cues with precision. In this study, we introduce a novel multi-modal VAD architecture, referred to as the Joint Actor-centric Visual, Audio, Language Encoder (JoVALE). JoVALE is the first VAD method to integrate audio and visual features with scene descriptive context sourced from large-capacity image captioning models. At the heart of JoVALE is the actor-centric aggregation of audio, visual, and scene descriptive information, enabling adaptive integration of crucial features for recognizing each actor's actions. We have developed a Transformer-based architecture, the Actor-centric Multi-modal Fusion Network, specifically designed to capture the dynamic interactions among actors and their multi-modal contexts. Our evaluation on three prominent VAD benchmarks, including AVA, UCF101-24, and JHMDB51-21, demonstrates that incorporating multi-modal information significantly enhances performance, setting new state-of-the-art performances in the field.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13708
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle JoVALE: Detecting Human Actions in Video Using Audiovisual and Language Contexts
Son, Taein
Seo, Soo Won
Kim, Jisong
Lee, Seok Hwan
Choi, Jun Won
Computer Vision and Pattern Recognition
Video Action Detection (VAD) entails localizing and categorizing action instances within videos, which inherently consist of diverse information sources such as audio, visual cues, and surrounding scene contexts. Leveraging this multi-modal information effectively for VAD poses a significant challenge, as the model must identify action-relevant cues with precision. In this study, we introduce a novel multi-modal VAD architecture, referred to as the Joint Actor-centric Visual, Audio, Language Encoder (JoVALE). JoVALE is the first VAD method to integrate audio and visual features with scene descriptive context sourced from large-capacity image captioning models. At the heart of JoVALE is the actor-centric aggregation of audio, visual, and scene descriptive information, enabling adaptive integration of crucial features for recognizing each actor's actions. We have developed a Transformer-based architecture, the Actor-centric Multi-modal Fusion Network, specifically designed to capture the dynamic interactions among actors and their multi-modal contexts. Our evaluation on three prominent VAD benchmarks, including AVA, UCF101-24, and JHMDB51-21, demonstrates that incorporating multi-modal information significantly enhances performance, setting new state-of-the-art performances in the field.
title JoVALE: Detecting Human Actions in Video Using Audiovisual and Language Contexts
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.13708