MS-CLR: Multi-Skeleton Contrastive Learning for Human Action Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kiray, Mert, Ritter, Alvaro, Navab, Nassir, Busam, Benjamin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915453101670400
author Kiray, Mert
Ritter, Alvaro
Navab, Nassir
Busam, Benjamin
author_facet Kiray, Mert
Ritter, Alvaro
Navab, Nassir
Busam, Benjamin
contents Contrastive learning has gained significant attention in skeleton-based action recognition for its ability to learn robust representations from unlabeled data. However, existing methods rely on a single skeleton convention, which limits their ability to generalize across datasets with diverse joint structures and anatomical coverage. We propose Multi-Skeleton Contrastive Learning (MS-CLR), a general self-supervised framework that aligns pose representations across multiple skeleton conventions extracted from the same sequence. This encourages the model to learn structural invariances and capture diverse anatomical cues, resulting in more expressive and generalizable features. To support this, we adapt the ST-GCN architecture to handle skeletons with varying joint layouts and scales through a unified representation scheme. Experiments on the NTU RGB+D 60 and 120 datasets demonstrate that MS-CLR consistently improves performance over strong single-skeleton contrastive learning baselines. A multi-skeleton ensemble further boosts performance, setting new state-of-the-art results on both datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2508_14889
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MS-CLR: Multi-Skeleton Contrastive Learning for Human Action Recognition
Kiray, Mert
Ritter, Alvaro
Navab, Nassir
Busam, Benjamin
Computer Vision and Pattern Recognition
Contrastive learning has gained significant attention in skeleton-based action recognition for its ability to learn robust representations from unlabeled data. However, existing methods rely on a single skeleton convention, which limits their ability to generalize across datasets with diverse joint structures and anatomical coverage. We propose Multi-Skeleton Contrastive Learning (MS-CLR), a general self-supervised framework that aligns pose representations across multiple skeleton conventions extracted from the same sequence. This encourages the model to learn structural invariances and capture diverse anatomical cues, resulting in more expressive and generalizable features. To support this, we adapt the ST-GCN architecture to handle skeletons with varying joint layouts and scales through a unified representation scheme. Experiments on the NTU RGB+D 60 and 120 datasets demonstrate that MS-CLR consistently improves performance over strong single-skeleton contrastive learning baselines. A multi-skeleton ensemble further boosts performance, setting new state-of-the-art results on both datasets.
title MS-CLR: Multi-Skeleton Contrastive Learning for Human Action Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.14889