Exploiting Ensemble Learning for Cross-View Isolated Sign Language Recognition

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Fei, Li, Kun, Nie, Yiqi, Duan, Zhangling, Zou, Peng, Wu, Zhiliang, Wang, Yuwei, Wei, Yanyan
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913677919125504
author Wang, Fei
Li, Kun
Nie, Yiqi
Duan, Zhangling
Zou, Peng
Wu, Zhiliang
Wang, Yuwei
Wei, Yanyan
author_facet Wang, Fei
Li, Kun
Nie, Yiqi
Duan, Zhangling
Zou, Peng
Wu, Zhiliang
Wang, Yuwei
Wei, Yanyan
contents In this paper, we present our solution to the Cross-View Isolated Sign Language Recognition (CV-ISLR) challenge held at WWW 2025. CV-ISLR addresses a critical issue in traditional Isolated Sign Language Recognition (ISLR), where existing datasets predominantly capture sign language videos from a frontal perspective, while real-world camera angles often vary. To accurately recognize sign language from different viewpoints, models must be capable of understanding gestures from multiple angles, making cross-view recognition challenging. To address this, we explore the advantages of ensemble learning, which enhances model robustness and generalization across diverse views. Our approach, built on a multi-dimensional Video Swin Transformer model, leverages this ensemble strategy to achieve competitive performance. Finally, our solution ranked 3rd in both the RGB-based ISLR and RGB-D-based ISLR tracks, demonstrating the effectiveness in handling the challenges of cross-view recognition. The code is available at: https://github.com/Jiafei127/CV_ISLR_WWW2025.
format Preprint
id arxiv_https___arxiv_org_abs_2502_02196
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploiting Ensemble Learning for Cross-View Isolated Sign Language Recognition
Wang, Fei
Li, Kun
Nie, Yiqi
Duan, Zhangling
Zou, Peng
Wu, Zhiliang
Wang, Yuwei
Wei, Yanyan
Computer Vision and Pattern Recognition
Artificial Intelligence
In this paper, we present our solution to the Cross-View Isolated Sign Language Recognition (CV-ISLR) challenge held at WWW 2025. CV-ISLR addresses a critical issue in traditional Isolated Sign Language Recognition (ISLR), where existing datasets predominantly capture sign language videos from a frontal perspective, while real-world camera angles often vary. To accurately recognize sign language from different viewpoints, models must be capable of understanding gestures from multiple angles, making cross-view recognition challenging. To address this, we explore the advantages of ensemble learning, which enhances model robustness and generalization across diverse views. Our approach, built on a multi-dimensional Video Swin Transformer model, leverages this ensemble strategy to achieve competitive performance. Finally, our solution ranked 3rd in both the RGB-based ISLR and RGB-D-based ISLR tracks, demonstrating the effectiveness in handling the challenges of cross-view recognition. The code is available at: https://github.com/Jiafei127/CV_ISLR_WWW2025.
title Exploiting Ensemble Learning for Cross-View Isolated Sign Language Recognition
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2502.02196