Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bao, Wentao, Li, Kai, Chen, Yuxiao, Patel, Deep, Min, Martin Renqiang, Kong, Yu
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913579765071872
author Bao, Wentao
Li, Kai
Chen, Yuxiao
Patel, Deep
Min, Martin Renqiang
Kong, Yu
author_facet Bao, Wentao
Li, Kai
Chen, Yuxiao
Patel, Deep
Min, Martin Renqiang
Kong, Yu
contents Action detection aims to detect (recognize and localize) human actions spatially and temporally in videos. Existing approaches focus on the closed-set setting where an action detector is trained and tested on videos from a fixed set of action categories. However, this constrained setting is not viable in an open world where test videos inevitably come beyond the trained action categories. In this paper, we address the practical yet challenging Open-Vocabulary Action Detection (OVAD) problem. It aims to detect any action in test videos while training a model on a fixed set of action categories. To achieve such an open-vocabulary capability, we propose a novel method OpenMixer that exploits the inherent semantics and localizability of large vision-language models (VLM) within the family of query-based detection transformers (DETR). Specifically, the OpenMixer is developed by spatial and temporal OpenMixer blocks (S-OMB and T-OMB), and a dynamically fused alignment (DFA) module. The three components collectively enjoy the merits of strong generalization from pre-trained VLMs and end-to-end learning from DETR design. Moreover, we established OVAD benchmarks under various settings, and the experimental results show that the OpenMixer performs the best over baselines for detecting seen and unseen actions. We release the codes, models, and dataset splits at https://github.com/Cogito2012/OpenMixer.
format Preprint
id arxiv_https___arxiv_org_abs_2411_10922
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection
Bao, Wentao
Li, Kai
Chen, Yuxiao
Patel, Deep
Min, Martin Renqiang
Kong, Yu
Computer Vision and Pattern Recognition
Action detection aims to detect (recognize and localize) human actions spatially and temporally in videos. Existing approaches focus on the closed-set setting where an action detector is trained and tested on videos from a fixed set of action categories. However, this constrained setting is not viable in an open world where test videos inevitably come beyond the trained action categories. In this paper, we address the practical yet challenging Open-Vocabulary Action Detection (OVAD) problem. It aims to detect any action in test videos while training a model on a fixed set of action categories. To achieve such an open-vocabulary capability, we propose a novel method OpenMixer that exploits the inherent semantics and localizability of large vision-language models (VLM) within the family of query-based detection transformers (DETR). Specifically, the OpenMixer is developed by spatial and temporal OpenMixer blocks (S-OMB and T-OMB), and a dynamically fused alignment (DFA) module. The three components collectively enjoy the merits of strong generalization from pre-trained VLMs and end-to-end learning from DETR design. Moreover, we established OVAD benchmarks under various settings, and the experimental results show that the OpenMixer performs the best over baselines for detecting seen and unseen actions. We release the codes, models, and dataset splits at https://github.com/Cogito2012/OpenMixer.
title Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.10922