SSL-SLR: Self-Supervised Representation Learning for Sign Language Recognition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Madjoukeng, Ariel Basso, Fink, Jérôme, Poitier, Pierre, Kenmogne, Edith Belise, Frenay, Benoit
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914373910396928
author Madjoukeng, Ariel Basso
Fink, Jérôme
Poitier, Pierre
Kenmogne, Edith Belise
Frenay, Benoit
author_facet Madjoukeng, Ariel Basso
Fink, Jérôme
Poitier, Pierre
Kenmogne, Edith Belise
Frenay, Benoit
contents Sign language recognition (SLR) is a machine learning task aiming to identify signs in videos. Due to the scarcity of annotated data, unsupervised methods like contrastive learning have become promising in this field. They learn meaningful representations by pulling positive pairs (two augmented versions of the same instance) closer and pushing negative pairs (different from the positive pairs) apart. In SLR, in a sign video, only certain parts provide information that is truly useful for its recognition. Applying contrastive methods to SLR raises two issues: (i) contrastive learning methods treat all parts of a video in the same way, without taking into account the relevance of certain parts over others; (ii) shared movements between different signs make negative pairs highly similar, complicating sign discrimination. These issues lead to learning non-discriminative features for sign recognition and poor results in downstream tasks. In response, this paper proposes a self-supervised learning framework designed to learn meaningful representations for SLR. This framework consists of two key components designed to work together: (i) a new self-supervised approach with free-negative pairs; (ii) a new data augmentation technique. This approach shows a considerable gain in accuracy compared to several contrastive and self-supervised methods, across linear evaluation, semi-supervised learning, and transferability between sign languages.
format Preprint
id arxiv_https___arxiv_org_abs_2509_05188
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SSL-SLR: Self-Supervised Representation Learning for Sign Language Recognition
Madjoukeng, Ariel Basso
Fink, Jérôme
Poitier, Pierre
Kenmogne, Edith Belise
Frenay, Benoit
Computer Vision and Pattern Recognition
Sign language recognition (SLR) is a machine learning task aiming to identify signs in videos. Due to the scarcity of annotated data, unsupervised methods like contrastive learning have become promising in this field. They learn meaningful representations by pulling positive pairs (two augmented versions of the same instance) closer and pushing negative pairs (different from the positive pairs) apart. In SLR, in a sign video, only certain parts provide information that is truly useful for its recognition. Applying contrastive methods to SLR raises two issues: (i) contrastive learning methods treat all parts of a video in the same way, without taking into account the relevance of certain parts over others; (ii) shared movements between different signs make negative pairs highly similar, complicating sign discrimination. These issues lead to learning non-discriminative features for sign recognition and poor results in downstream tasks. In response, this paper proposes a self-supervised learning framework designed to learn meaningful representations for SLR. This framework consists of two key components designed to work together: (i) a new self-supervised approach with free-negative pairs; (ii) a new data augmentation technique. This approach shows a considerable gain in accuracy compared to several contrastive and self-supervised methods, across linear evaluation, semi-supervised learning, and transferability between sign languages.
title SSL-SLR: Self-Supervised Representation Learning for Sign Language Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.05188