Spatio-Temporal Joint Density Driven Learning for Skeleton-Based Action Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gunasekara, Shanaka Ramesh, Li, Wanqing, Ogunbona, Philip, Yang, Jack
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912401010458624
author Gunasekara, Shanaka Ramesh
Li, Wanqing
Ogunbona, Philip
Yang, Jack
author_facet Gunasekara, Shanaka Ramesh
Li, Wanqing
Ogunbona, Philip
Yang, Jack
contents Traditional approaches in unsupervised or self supervised learning for skeleton-based action classification have concentrated predominantly on the dynamic aspects of skeletal sequences. Yet, the intricate interaction between the moving and static elements of the skeleton presents a rarely tapped discriminative potential for action classification. This paper introduces a novel measurement, referred to as spatial-temporal joint density (STJD), to quantify such interaction. Tracking the evolution of this density throughout an action can effectively identify a subset of discriminative moving and/or static joints termed "prime joints" to steer self-supervised learning. A new contrastive learning strategy named STJD-CL is proposed to align the representation of a skeleton sequence with that of its prime joints while simultaneously contrasting the representations of prime and nonprime joints. In addition, a method called STJD-MP is developed by integrating it with a reconstruction-based framework for more effective learning. Experimental evaluations on the NTU RGB+D 60, NTU RGB+D 120, and PKUMMD datasets in various downstream tasks demonstrate that the proposed STJD-CL and STJD-MP improved performance, particularly by 3.5 and 3.6 percentage points over the state-of-the-art contrastive methods on the NTU RGB+D 120 dataset using X-sub and X-set evaluations, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23012
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Spatio-Temporal Joint Density Driven Learning for Skeleton-Based Action Recognition
Gunasekara, Shanaka Ramesh
Li, Wanqing
Ogunbona, Philip
Yang, Jack
Computer Vision and Pattern Recognition
Traditional approaches in unsupervised or self supervised learning for skeleton-based action classification have concentrated predominantly on the dynamic aspects of skeletal sequences. Yet, the intricate interaction between the moving and static elements of the skeleton presents a rarely tapped discriminative potential for action classification. This paper introduces a novel measurement, referred to as spatial-temporal joint density (STJD), to quantify such interaction. Tracking the evolution of this density throughout an action can effectively identify a subset of discriminative moving and/or static joints termed "prime joints" to steer self-supervised learning. A new contrastive learning strategy named STJD-CL is proposed to align the representation of a skeleton sequence with that of its prime joints while simultaneously contrasting the representations of prime and nonprime joints. In addition, a method called STJD-MP is developed by integrating it with a reconstruction-based framework for more effective learning. Experimental evaluations on the NTU RGB+D 60, NTU RGB+D 120, and PKUMMD datasets in various downstream tasks demonstrate that the proposed STJD-CL and STJD-MP improved performance, particularly by 3.5 and 3.6 percentage points over the state-of-the-art contrastive methods on the NTU RGB+D 120 dataset using X-sub and X-set evaluations, respectively.
title Spatio-Temporal Joint Density Driven Learning for Skeleton-Based Action Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.23012