Hierarchical Action Recognition: A Contrastive Video-Language Approach with Hierarchical Interactions

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhang, Rui, Li, Shuailong, Xue, Junxiao, Lin, Feng, Zhang, Qing, Ma, Xiao, Yan, Xiaoran
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910460588064768
author Zhang, Rui
Li, Shuailong
Xue, Junxiao
Lin, Feng
Zhang, Qing
Ma, Xiao
Yan, Xiaoran
author_facet Zhang, Rui
Li, Shuailong
Xue, Junxiao
Lin, Feng
Zhang, Qing
Ma, Xiao
Yan, Xiaoran
contents Video recognition remains an open challenge, requiring the identification of diverse content categories within videos. Mainstream approaches often perform flat classification, overlooking the intrinsic hierarchical structure relating categories. To address this, we formalize the novel task of hierarchical video recognition, and propose a video-language learning framework tailored for hierarchical recognition. Specifically, our framework encodes dependencies between hierarchical category levels, and applies a top-down constraint to filter recognition predictions. We further construct a new fine-grained dataset based on medical assessments for rehabilitation of stroke patients, serving as a challenging benchmark for hierarchical recognition. Through extensive experiments, we demonstrate the efficacy of our approach for hierarchical recognition, significantly outperforming conventional methods, especially for fine-grained subcategories. The proposed framework paves the way for hierarchical modeling in video understanding tasks, moving beyond flat categorization.
format Preprint
id arxiv_https___arxiv_org_abs_2405_17729
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hierarchical Action Recognition: A Contrastive Video-Language Approach with Hierarchical Interactions
Zhang, Rui
Li, Shuailong
Xue, Junxiao
Lin, Feng
Zhang, Qing
Ma, Xiao
Yan, Xiaoran
Computer Vision and Pattern Recognition
Multimedia
Video recognition remains an open challenge, requiring the identification of diverse content categories within videos. Mainstream approaches often perform flat classification, overlooking the intrinsic hierarchical structure relating categories. To address this, we formalize the novel task of hierarchical video recognition, and propose a video-language learning framework tailored for hierarchical recognition. Specifically, our framework encodes dependencies between hierarchical category levels, and applies a top-down constraint to filter recognition predictions. We further construct a new fine-grained dataset based on medical assessments for rehabilitation of stroke patients, serving as a challenging benchmark for hierarchical recognition. Through extensive experiments, we demonstrate the efficacy of our approach for hierarchical recognition, significantly outperforming conventional methods, especially for fine-grained subcategories. The proposed framework paves the way for hierarchical modeling in video understanding tasks, moving beyond flat categorization.
title Hierarchical Action Recognition: A Contrastive Video-Language Approach with Hierarchical Interactions
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2405.17729