Isharah: A Large-Scale Multi-Scene Dataset for Continuous Sign Language Recognition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Alyami, Sarah, Luqman, Hamzah, Al-Azani, Sadam, Alowaifeer, Maad, Alharbi, Yazeed, Alonaizan, Yaser
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915324248457216
author Alyami, Sarah
Luqman, Hamzah
Al-Azani, Sadam
Alowaifeer, Maad
Alharbi, Yazeed
Alonaizan, Yaser
author_facet Alyami, Sarah
Luqman, Hamzah
Al-Azani, Sadam
Alowaifeer, Maad
Alharbi, Yazeed
Alonaizan, Yaser
contents Current benchmarks for sign language recognition (SLR) focus mainly on isolated SLR, while there are limited datasets for continuous SLR (CSLR), which recognizes sequences of signs in a video. Additionally, existing CSLR datasets are collected in controlled settings, which restricts their effectiveness in building robust real-world CSLR systems. To address these limitations, we present Isharah, a large multi-scene dataset for CSLR. It is the first dataset of its type and size that has been collected in an unconstrained environment using signers' smartphone cameras. This setup resulted in high variations of recording settings, camera distances, angles, and resolutions. This variation helps with developing sign language understanding models capable of handling the variability and complexity of real-world scenarios. The dataset consists of 30,000 video clips performed by 18 deaf and professional signers. Additionally, the dataset is linguistically rich as it provides a gloss-level annotation for all dataset's videos, making it useful for developing CSLR and sign language translation (SLT) systems. This paper also introduces multiple sign language understanding benchmarks, including signer-independent and unseen-sentence CSLR, along with gloss-based and gloss-free SLT. The Isharah dataset is available on https://snalyami.github.io/Isharah_CSLR/.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03615
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Isharah: A Large-Scale Multi-Scene Dataset for Continuous Sign Language Recognition
Alyami, Sarah
Luqman, Hamzah
Al-Azani, Sadam
Alowaifeer, Maad
Alharbi, Yazeed
Alonaizan, Yaser
Computer Vision and Pattern Recognition
Current benchmarks for sign language recognition (SLR) focus mainly on isolated SLR, while there are limited datasets for continuous SLR (CSLR), which recognizes sequences of signs in a video. Additionally, existing CSLR datasets are collected in controlled settings, which restricts their effectiveness in building robust real-world CSLR systems. To address these limitations, we present Isharah, a large multi-scene dataset for CSLR. It is the first dataset of its type and size that has been collected in an unconstrained environment using signers' smartphone cameras. This setup resulted in high variations of recording settings, camera distances, angles, and resolutions. This variation helps with developing sign language understanding models capable of handling the variability and complexity of real-world scenarios. The dataset consists of 30,000 video clips performed by 18 deaf and professional signers. Additionally, the dataset is linguistically rich as it provides a gloss-level annotation for all dataset's videos, making it useful for developing CSLR and sign language translation (SLT) systems. This paper also introduces multiple sign language understanding benchmarks, including signer-independent and unseen-sentence CSLR, along with gloss-based and gloss-free SLT. The Isharah dataset is available on https://snalyami.github.io/Isharah_CSLR/.
title Isharah: A Large-Scale Multi-Scene Dataset for Continuous Sign Language Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.03615