P2ANet: A Dataset and Benchmark for Dense Action Detection from Table Tennis Match Broadcasting Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bian, Jiang, Li, Xuhong, Wang, Tao, Wang, Qingzhong, Huang, Jun, Liu, Chen, Zhao, Jun, Lu, Feixiang, Dou, Dejing, Xiong, Haoyi
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914728524120064
author Bian, Jiang
Li, Xuhong
Wang, Tao
Wang, Qingzhong
Huang, Jun
Liu, Chen
Zhao, Jun
Lu, Feixiang
Dou, Dejing
Xiong, Haoyi
author_facet Bian, Jiang
Li, Xuhong
Wang, Tao
Wang, Qingzhong
Huang, Jun
Liu, Chen
Zhao, Jun
Lu, Feixiang
Dou, Dejing
Xiong, Haoyi
contents While deep learning has been widely used for video analytics, such as video classification and action detection, dense action detection with fast-moving subjects from sports videos is still challenging. In this work, we release yet another sports video benchmark \TheName{} for \emph{\underline{P}}ing \emph{\underline{P}}ong-\emph{\underline{A}}ction detection, which consists of 2,721 video clips collected from the broadcasting videos of professional table tennis matches in World Table Tennis Championships and Olympiads. We work with a crew of table tennis professionals and referees on a specially designed annotation toolbox to obtain fine-grained action labels (in 14 classes) for every ping-pong action that appeared in the dataset, and formulate two sets of action detection problems -- \emph{action localization} and \emph{action recognition}. We evaluate a number of commonly-seen action recognition (e.g., TSM, TSN, Video SwinTransformer, and Slowfast) and action localization models (e.g., BSN, BSN++, BMN, TCANet), using \TheName{} for both problems, under various settings. These models can only achieve 48\% area under the AR-AN curve for localization and 82\% top-one accuracy for recognition since the ping-pong actions are dense with fast-moving subjects but broadcasting videos are with only 25 FPS. The results confirm that \TheName{} is still a challenging task and can be used as a special benchmark for dense action detection from videos.
format Preprint
id arxiv_https___arxiv_org_abs_2207_12730
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle P2ANet: A Dataset and Benchmark for Dense Action Detection from Table Tennis Match Broadcasting Videos
Bian, Jiang
Li, Xuhong
Wang, Tao
Wang, Qingzhong
Huang, Jun
Liu, Chen
Zhao, Jun
Lu, Feixiang
Dou, Dejing
Xiong, Haoyi
Computer Vision and Pattern Recognition
Machine Learning
While deep learning has been widely used for video analytics, such as video classification and action detection, dense action detection with fast-moving subjects from sports videos is still challenging. In this work, we release yet another sports video benchmark \TheName{} for \emph{\underline{P}}ing \emph{\underline{P}}ong-\emph{\underline{A}}ction detection, which consists of 2,721 video clips collected from the broadcasting videos of professional table tennis matches in World Table Tennis Championships and Olympiads. We work with a crew of table tennis professionals and referees on a specially designed annotation toolbox to obtain fine-grained action labels (in 14 classes) for every ping-pong action that appeared in the dataset, and formulate two sets of action detection problems -- \emph{action localization} and \emph{action recognition}. We evaluate a number of commonly-seen action recognition (e.g., TSM, TSN, Video SwinTransformer, and Slowfast) and action localization models (e.g., BSN, BSN++, BMN, TCANet), using \TheName{} for both problems, under various settings. These models can only achieve 48\% area under the AR-AN curve for localization and 82\% top-one accuracy for recognition since the ping-pong actions are dense with fast-moving subjects but broadcasting videos are with only 25 FPS. The results confirm that \TheName{} is still a challenging task and can be used as a special benchmark for dense action detection from videos.
title P2ANet: A Dataset and Benchmark for Dense Action Detection from Table Tennis Match Broadcasting Videos
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2207.12730