Action Segmentation Using 2D Skeleton Heatmaps and Multi-Modality Fusion

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Hyder, Syed Waleed, Usama, Muhammad, Zafar, Anas, Naufil, Muhammad, Fateh, Fawad Javed, Konin, Andrey, Zia, M. Zeeshan, Tran, Quoc-Huy
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909182493458432
author Hyder, Syed Waleed
Usama, Muhammad
Zafar, Anas
Naufil, Muhammad
Fateh, Fawad Javed
Konin, Andrey
Zia, M. Zeeshan
Tran, Quoc-Huy
author_facet Hyder, Syed Waleed
Usama, Muhammad
Zafar, Anas
Naufil, Muhammad
Fateh, Fawad Javed
Konin, Andrey
Zia, M. Zeeshan
Tran, Quoc-Huy
contents This paper presents a 2D skeleton-based action segmentation method with applications in fine-grained human activity recognition. In contrast with state-of-the-art methods which directly take sequences of 3D skeleton coordinates as inputs and apply Graph Convolutional Networks (GCNs) for spatiotemporal feature learning, our main idea is to use sequences of 2D skeleton heatmaps as inputs and employ Temporal Convolutional Networks (TCNs) to extract spatiotemporal features. Despite lacking 3D information, our approach yields comparable/superior performances and better robustness against missing keypoints than previous methods on action segmentation datasets. Moreover, we improve the performances further by using both 2D skeleton heatmaps and RGB videos as inputs. To our best knowledge, this is the first work to utilize 2D skeleton heatmap inputs and the first work to explore 2D skeleton+RGB fusion for action segmentation.
format Preprint
id arxiv_https___arxiv_org_abs_2309_06462
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Action Segmentation Using 2D Skeleton Heatmaps and Multi-Modality Fusion
Hyder, Syed Waleed
Usama, Muhammad
Zafar, Anas
Naufil, Muhammad
Fateh, Fawad Javed
Konin, Andrey
Zia, M. Zeeshan
Tran, Quoc-Huy
Computer Vision and Pattern Recognition
This paper presents a 2D skeleton-based action segmentation method with applications in fine-grained human activity recognition. In contrast with state-of-the-art methods which directly take sequences of 3D skeleton coordinates as inputs and apply Graph Convolutional Networks (GCNs) for spatiotemporal feature learning, our main idea is to use sequences of 2D skeleton heatmaps as inputs and employ Temporal Convolutional Networks (TCNs) to extract spatiotemporal features. Despite lacking 3D information, our approach yields comparable/superior performances and better robustness against missing keypoints than previous methods on action segmentation datasets. Moreover, we improve the performances further by using both 2D skeleton heatmaps and RGB videos as inputs. To our best knowledge, this is the first work to utilize 2D skeleton heatmap inputs and the first work to explore 2D skeleton+RGB fusion for action segmentation.
title Action Segmentation Using 2D Skeleton Heatmaps and Multi-Modality Fusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2309.06462