American Sign Language Video to Text Translation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909102169391104 |
|---|---|
| author | Roy, Parsheeta Han, Ji-Eun Chouhan, Srishti Thumu, Bhaavanaa |
| author_facet | Roy, Parsheeta Han, Ji-Eun Chouhan, Srishti Thumu, Bhaavanaa |
| contents | Sign language to text is a crucial technology that can break down communication barriers for individuals with hearing difficulties. We replicate and try to improve on a recently published study. We evaluate models using BLEU and rBLEU metrics to ensure translation quality. During our ablation study, we found that the model's performance is significantly influenced by optimizers, activation functions, and label smoothing. Further research aims to refine visual feature capturing, enhance decoder utilization, and integrate pre-trained decoders for better translation outcomes. Our source code is available to facilitate replication of our results and encourage future research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_07255 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | American Sign Language Video to Text Translation Roy, Parsheeta Han, Ji-Eun Chouhan, Srishti Thumu, Bhaavanaa Computation and Language Computer Vision and Pattern Recognition Sign language to text is a crucial technology that can break down communication barriers for individuals with hearing difficulties. We replicate and try to improve on a recently published study. We evaluate models using BLEU and rBLEU metrics to ensure translation quality. During our ablation study, we found that the model's performance is significantly influenced by optimizers, activation functions, and label smoothing. Further research aims to refine visual feature capturing, enhance decoder utilization, and integrate pre-trained decoders for better translation outcomes. Our source code is available to facilitate replication of our results and encourage future research. |
| title | American Sign Language Video to Text Translation |
| topic | Computation and Language Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2402.07255 |