Sign Language Translation using Frame and Event Stream: Benchmark Dataset and Algorithms

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xiao, Li, Yuehang, Wang, Fuling, Jiang, Bo, Wang, Yaowei, Tian, Yonghong, Tang, Jin, Luo, Bin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910865346789376
author Wang, Xiao
Li, Yuehang
Wang, Fuling
Jiang, Bo
Wang, Yaowei
Tian, Yonghong
Tang, Jin
Luo, Bin
author_facet Wang, Xiao
Li, Yuehang
Wang, Fuling
Jiang, Bo
Wang, Yaowei
Tian, Yonghong
Tang, Jin
Luo, Bin
contents Accurate sign language understanding serves as a crucial communication channel for individuals with disabilities. Current sign language translation algorithms predominantly rely on RGB frames, which may be limited by fixed frame rates, variable lighting conditions, and motion blur caused by rapid hand movements. Inspired by the recent successful application of event cameras in other fields, we propose to leverage event streams to assist RGB cameras in capturing gesture data, addressing the various challenges mentioned above. Specifically, we first collect a large-scale RGB-Event sign language translation dataset using the DVS346 camera, termed VECSL, which contains 15,676 RGB-Event samples, 15,191 glosses, and covers 2,568 Chinese characters. These samples were gathered across a diverse range of indoor and outdoor environments, capturing multiple viewing angles, varying light intensities, and different camera motions. Due to the absence of benchmark algorithms for comparison in this new task, we retrained and evaluated multiple state-of-the-art SLT algorithms, and believe that this benchmark can effectively support subsequent related research. Additionally, we propose a novel RGB-Event sign language translation framework (i.e., M$^2$-SLT) that incorporates fine-grained micro-sign and coarse-grained macro-sign retrieval, achieving state-of-the-art results on the proposed dataset. Both the source code and dataset will be released on https://github.com/Event-AHU/OpenESL.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06484
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sign Language Translation using Frame and Event Stream: Benchmark Dataset and Algorithms
Wang, Xiao
Li, Yuehang
Wang, Fuling
Jiang, Bo
Wang, Yaowei
Tian, Yonghong
Tang, Jin
Luo, Bin
Computer Vision and Pattern Recognition
Artificial Intelligence
Neural and Evolutionary Computing
Accurate sign language understanding serves as a crucial communication channel for individuals with disabilities. Current sign language translation algorithms predominantly rely on RGB frames, which may be limited by fixed frame rates, variable lighting conditions, and motion blur caused by rapid hand movements. Inspired by the recent successful application of event cameras in other fields, we propose to leverage event streams to assist RGB cameras in capturing gesture data, addressing the various challenges mentioned above. Specifically, we first collect a large-scale RGB-Event sign language translation dataset using the DVS346 camera, termed VECSL, which contains 15,676 RGB-Event samples, 15,191 glosses, and covers 2,568 Chinese characters. These samples were gathered across a diverse range of indoor and outdoor environments, capturing multiple viewing angles, varying light intensities, and different camera motions. Due to the absence of benchmark algorithms for comparison in this new task, we retrained and evaluated multiple state-of-the-art SLT algorithms, and believe that this benchmark can effectively support subsequent related research. Additionally, we propose a novel RGB-Event sign language translation framework (i.e., M$^2$-SLT) that incorporates fine-grained micro-sign and coarse-grained macro-sign retrieval, achieving state-of-the-art results on the proposed dataset. Both the source code and dataset will be released on https://github.com/Event-AHU/OpenESL.
title Sign Language Translation using Frame and Event Stream: Benchmark Dataset and Algorithms
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Neural and Evolutionary Computing
url https://arxiv.org/abs/2503.06484