Bangla Sign Language Translation: Dataset Creation Challenges, Benchmarking and Prospects

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rubaiyeat, Husne Ara, Mahmud, Hasan, Hasan, Md Kamrul
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914171925299200
author Rubaiyeat, Husne Ara
Mahmud, Hasan
Hasan, Md Kamrul
author_facet Rubaiyeat, Husne Ara
Mahmud, Hasan
Hasan, Md Kamrul
contents Bangla Sign Language Translation (BdSLT) has been severely constrained so far as the language itself is very low resource. Standard sentence level dataset creation for BdSLT is of immense importance for developing AI based assistive tools for deaf and hard of hearing people of Bangla speaking community. In this paper, we present a dataset, IsharaKhobor , and two subset of it for enabling research. We also present the challenges towards developing the dataset and present some way forward by benchmarking with landmark based raw and RQE embedding. We do some ablation on vocabulary restriction and canonicalization of the same within the dataset, which resulted in two more datasets, IsharaKhobor_small and IsharaKhobor_canonical_small. The dataset is publicly available at: www.kaggle.com/datasets/hasanssl/isharakhobor [1].
format Preprint
id arxiv_https___arxiv_org_abs_2511_21533
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bangla Sign Language Translation: Dataset Creation Challenges, Benchmarking and Prospects
Rubaiyeat, Husne Ara
Mahmud, Hasan
Hasan, Md Kamrul
Computation and Language
Computer Vision and Pattern Recognition
Bangla Sign Language Translation (BdSLT) has been severely constrained so far as the language itself is very low resource. Standard sentence level dataset creation for BdSLT is of immense importance for developing AI based assistive tools for deaf and hard of hearing people of Bangla speaking community. In this paper, we present a dataset, IsharaKhobor , and two subset of it for enabling research. We also present the challenges towards developing the dataset and present some way forward by benchmarking with landmark based raw and RQE embedding. We do some ablation on vocabulary restriction and canonicalization of the same within the dataset, which resulted in two more datasets, IsharaKhobor_small and IsharaKhobor_canonical_small. The dataset is publicly available at: www.kaggle.com/datasets/hasanssl/isharakhobor [1].
title Bangla Sign Language Translation: Dataset Creation Challenges, Benchmarking and Prospects
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.21533