ContextDet: Temporal Action Detection with Adaptive Context Aggregation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Ning, Xiao, Yun, Peng, Xiaopeng, Chang, Xiaojun, Wang, Xuanhong, Fang, Dingyi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917809993285632
author Wang, Ning
Xiao, Yun
Peng, Xiaopeng
Chang, Xiaojun
Wang, Xuanhong
Fang, Dingyi
author_facet Wang, Ning
Xiao, Yun
Peng, Xiaopeng
Chang, Xiaojun
Wang, Xuanhong
Fang, Dingyi
contents Temporal action detection (TAD), which locates and recognizes action segments, remains a challenging task in video understanding due to variable segment lengths and ambiguous boundaries. Existing methods treat neighboring contexts of an action segment indiscriminately, leading to imprecise boundary predictions. We introduce a single-stage ContextDet framework, which makes use of large-kernel convolutions in TAD for the first time. Our model features a pyramid adaptive context aggragation (ACA) architecture, capturing long context and improving action discriminability. Each ACA level consists of two novel modules. The context attention module (CAM) identifies salient contextual information, encourages context diversity, and preserves context integrity through a context gating block (CGB). The long context module (LCM) makes use of a mixture of large- and small-kernel convolutions to adaptively gather long-range context and fine-grained local features. Additionally, by varying the length of these large kernels across the ACA pyramid, our model provides lightweight yet effective context aggregation and action discrimination. We conducted extensive experiments and compared our model with a number of advanced TAD methods on six challenging TAD benchmarks: MultiThumos, Charades, FineAction, EPIC-Kitchens 100, Thumos14, and HACS, demonstrating superior accuracy at reduced inference speed.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15279
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ContextDet: Temporal Action Detection with Adaptive Context Aggregation
Wang, Ning
Xiao, Yun
Peng, Xiaopeng
Chang, Xiaojun
Wang, Xuanhong
Fang, Dingyi
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
Temporal action detection (TAD), which locates and recognizes action segments, remains a challenging task in video understanding due to variable segment lengths and ambiguous boundaries. Existing methods treat neighboring contexts of an action segment indiscriminately, leading to imprecise boundary predictions. We introduce a single-stage ContextDet framework, which makes use of large-kernel convolutions in TAD for the first time. Our model features a pyramid adaptive context aggragation (ACA) architecture, capturing long context and improving action discriminability. Each ACA level consists of two novel modules. The context attention module (CAM) identifies salient contextual information, encourages context diversity, and preserves context integrity through a context gating block (CGB). The long context module (LCM) makes use of a mixture of large- and small-kernel convolutions to adaptively gather long-range context and fine-grained local features. Additionally, by varying the length of these large kernels across the ACA pyramid, our model provides lightweight yet effective context aggregation and action discrimination. We conducted extensive experiments and compared our model with a number of advanced TAD methods on six challenging TAD benchmarks: MultiThumos, Charades, FineAction, EPIC-Kitchens 100, Thumos14, and HACS, demonstrating superior accuracy at reduced inference speed.
title ContextDet: Temporal Action Detection with Adaptive Context Aggregation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2410.15279