American Sign Language Video to Text Translation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Roy, Parsheeta, Han, Ji-Eun, Chouhan, Srishti, Thumu, Bhaavanaa
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909102169391104
author Roy, Parsheeta
Han, Ji-Eun
Chouhan, Srishti
Thumu, Bhaavanaa
author_facet Roy, Parsheeta
Han, Ji-Eun
Chouhan, Srishti
Thumu, Bhaavanaa
contents Sign language to text is a crucial technology that can break down communication barriers for individuals with hearing difficulties. We replicate and try to improve on a recently published study. We evaluate models using BLEU and rBLEU metrics to ensure translation quality. During our ablation study, we found that the model's performance is significantly influenced by optimizers, activation functions, and label smoothing. Further research aims to refine visual feature capturing, enhance decoder utilization, and integrate pre-trained decoders for better translation outcomes. Our source code is available to facilitate replication of our results and encourage future research.
format Preprint
id arxiv_https___arxiv_org_abs_2402_07255
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle American Sign Language Video to Text Translation
Roy, Parsheeta
Han, Ji-Eun
Chouhan, Srishti
Thumu, Bhaavanaa
Computation and Language
Computer Vision and Pattern Recognition
Sign language to text is a crucial technology that can break down communication barriers for individuals with hearing difficulties. We replicate and try to improve on a recently published study. We evaluate models using BLEU and rBLEU metrics to ensure translation quality. During our ablation study, we found that the model's performance is significantly influenced by optimizers, activation functions, and label smoothing. Further research aims to refine visual feature capturing, enhance decoder utilization, and integrate pre-trained decoders for better translation outcomes. Our source code is available to facilitate replication of our results and encourage future research.
title American Sign Language Video to Text Translation
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.07255