Cross-Modal Consistency Learning for Sign Language Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Kepeng, Li, Zecheng, Hu, Hezhen, Zhou, Wengang, Li, Houqiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909546882007040
author Wu, Kepeng
Li, Zecheng
Hu, Hezhen
Zhou, Wengang
Li, Houqiang
author_facet Wu, Kepeng
Li, Zecheng
Hu, Hezhen
Zhou, Wengang
Li, Houqiang
contents Pre-training has been proven to be effective in boosting the performance of Isolated Sign Language Recognition (ISLR). Existing pre-training methods solely focus on the compact pose data, which eliminates background perturbation but inevitably suffers from insufficient semantic cues compared to raw RGB videos. Nevertheless, learning representation directly from RGB videos remains challenging due to the presence of sign-independent visual features. To address this dilemma, we propose a Cross-modal Consistency Learning framework (CCL-SLR), which leverages the cross-modal consistency from both RGB and pose modalities based on self-supervised pre-training. First, CCL-SLR employs contrastive learning for instance discrimination within and across modalities. Through the single-modal and cross-modal contrastive learning, CCL-SLR gradually aligns the feature spaces of RGB and pose modalities, thereby extracting consistent sign representations. Second, we further introduce Motion-Preserving Masking (MPM) and Semantic Positive Mining (SPM) techniques to improve cross-modal consistency from the perspective of data augmentation and sample similarity, respectively. Extensive experiments on four ISLR benchmarks show that CCL-SLR achieves impressive performance, demonstrating its effectiveness. The code will be released to the public.
format Preprint
id arxiv_https___arxiv_org_abs_2503_12485
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Cross-Modal Consistency Learning for Sign Language Recognition
Wu, Kepeng
Li, Zecheng
Hu, Hezhen
Zhou, Wengang
Li, Houqiang
Computer Vision and Pattern Recognition
Pre-training has been proven to be effective in boosting the performance of Isolated Sign Language Recognition (ISLR). Existing pre-training methods solely focus on the compact pose data, which eliminates background perturbation but inevitably suffers from insufficient semantic cues compared to raw RGB videos. Nevertheless, learning representation directly from RGB videos remains challenging due to the presence of sign-independent visual features. To address this dilemma, we propose a Cross-modal Consistency Learning framework (CCL-SLR), which leverages the cross-modal consistency from both RGB and pose modalities based on self-supervised pre-training. First, CCL-SLR employs contrastive learning for instance discrimination within and across modalities. Through the single-modal and cross-modal contrastive learning, CCL-SLR gradually aligns the feature spaces of RGB and pose modalities, thereby extracting consistent sign representations. Second, we further introduce Motion-Preserving Masking (MPM) and Semantic Positive Mining (SPM) techniques to improve cross-modal consistency from the perspective of data augmentation and sample similarity, respectively. Extensive experiments on four ISLR benchmarks show that CCL-SLR achieves impressive performance, demonstrating its effectiveness. The code will be released to the public.
title Cross-Modal Consistency Learning for Sign Language Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.12485