Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zuo, Ronglai, Wei, Fangyun, Mak, Brian
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:https://arxiv.org/abs/2401.05336
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914954061283328
author Zuo, Ronglai
Wei, Fangyun
Mak, Brian
author_facet Zuo, Ronglai
Wei, Fangyun
Mak, Brian
contents Research on continuous sign language recognition (CSLR) is essential to bridge the communication gap between deaf and hearing individuals. Numerous previous studies have trained their models using the connectionist temporal classification (CTC) loss. During inference, these CTC-based models generally require the entire sign video as input to make predictions, a process known as offline recognition, which suffers from high latency and substantial memory usage. In this work, we take the first step towards online CSLR. Our approach consists of three phases: 1) developing a sign dictionary; 2) training an isolated sign language recognition model on the dictionary; and 3) employing a sliding window approach on the input sign sequence, feeding each sign clip to the optimized model for online recognition. Additionally, our online recognition model can be extended to support online translation by integrating a gloss-to-text network and can enhance the performance of any offline model. With these extensions, our online approach achieves new state-of-the-art performance on three popular benchmarks across various task settings. Code and models are available at https://github.com/FangyunWei/SLRT.
format Preprint
id arxiv_https___arxiv_org_abs_2401_05336
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Online Continuous Sign Language Recognition and Translation
Zuo, Ronglai
Wei, Fangyun
Mak, Brian
Computer Vision and Pattern Recognition
Research on continuous sign language recognition (CSLR) is essential to bridge the communication gap between deaf and hearing individuals. Numerous previous studies have trained their models using the connectionist temporal classification (CTC) loss. During inference, these CTC-based models generally require the entire sign video as input to make predictions, a process known as offline recognition, which suffers from high latency and substantial memory usage. In this work, we take the first step towards online CSLR. Our approach consists of three phases: 1) developing a sign dictionary; 2) training an isolated sign language recognition model on the dictionary; and 3) employing a sliding window approach on the input sign sequence, feeding each sign clip to the optimized model for online recognition. Additionally, our online recognition model can be extended to support online translation by integrating a gloss-to-text network and can enhance the performance of any offline model. With these extensions, our online approach achieves new state-of-the-art performance on three popular benchmarks across various task settings. Code and models are available at https://github.com/FangyunWei/SLRT.
title Towards Online Continuous Sign Language Recognition and Translation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.05336