Spatio-temporal Sign Language Representation and Translation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hamidullah, Yasser, van Genabith, Josef, España-Bonet, Cristina
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918165564358656
author Hamidullah, Yasser
van Genabith, Josef
España-Bonet, Cristina
author_facet Hamidullah, Yasser
van Genabith, Josef
España-Bonet, Cristina
contents This paper describes the DFKI-MLT submission to the WMT-SLT 2022 sign language translation (SLT) task from Swiss German Sign Language (video) into German (text). State-of-the-art techniques for SLT use a generic seq2seq architecture with customized input embeddings. Instead of word embeddings as used in textual machine translation, SLT systems use features extracted from video frames. Standard approaches often do not benefit from temporal features. In our participation, we present a system that learns spatio-temporal feature representations and translation in a single model, resulting in a real end-to-end architecture expected to better generalize to new data sets. Our best system achieved $5\pm1$ BLEU points on the development set, but the performance on the test dropped to $0.11\pm0.06$ BLEU points.
format Preprint
id arxiv_https___arxiv_org_abs_2510_19413
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Spatio-temporal Sign Language Representation and Translation
Hamidullah, Yasser
van Genabith, Josef
España-Bonet, Cristina
Computation and Language
Computer Vision and Pattern Recognition
This paper describes the DFKI-MLT submission to the WMT-SLT 2022 sign language translation (SLT) task from Swiss German Sign Language (video) into German (text). State-of-the-art techniques for SLT use a generic seq2seq architecture with customized input embeddings. Instead of word embeddings as used in textual machine translation, SLT systems use features extracted from video frames. Standard approaches often do not benefit from temporal features. In our participation, we present a system that learns spatio-temporal feature representations and translation in a single model, resulting in a real end-to-end architecture expected to better generalize to new data sets. Our best system achieved $5\pm1$ BLEU points on the development set, but the performance on the test dropped to $0.11\pm0.06$ BLEU points.
title Spatio-temporal Sign Language Representation and Translation
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.19413