Sign Language Translation with Sentence Embedding Supervision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hamidullah, Yasser, van Genabith, Josef, España-Bonet, Cristina
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911226607435776
author Hamidullah, Yasser
van Genabith, Josef
España-Bonet, Cristina
author_facet Hamidullah, Yasser
van Genabith, Josef
España-Bonet, Cristina
contents State-of-the-art sign language translation (SLT) systems facilitate the learning process through gloss annotations, either in an end2end manner or by involving an intermediate step. Unfortunately, gloss labelled sign language data is usually not available at scale and, when available, gloss annotations widely differ from dataset to dataset. We present a novel approach using sentence embeddings of the target sentences at training time that take the role of glosses. The new kind of supervision does not need any manual annotation but it is learned on raw textual data. As our approach easily facilitates multilinguality, we evaluate it on datasets covering German (PHOENIX-2014T) and American (How2Sign) sign languages and experiment with mono- and multilingual sentence embeddings and translation systems. Our approach significantly outperforms other gloss-free approaches, setting the new state-of-the-art for data sets where glosses are not available and when no additional SLT datasets are used for pretraining, diminishing the gap between gloss-free and gloss-dependent systems.
format Preprint
id arxiv_https___arxiv_org_abs_2510_19367
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sign Language Translation with Sentence Embedding Supervision
Hamidullah, Yasser
van Genabith, Josef
España-Bonet, Cristina
Computation and Language
State-of-the-art sign language translation (SLT) systems facilitate the learning process through gloss annotations, either in an end2end manner or by involving an intermediate step. Unfortunately, gloss labelled sign language data is usually not available at scale and, when available, gloss annotations widely differ from dataset to dataset. We present a novel approach using sentence embeddings of the target sentences at training time that take the role of glosses. The new kind of supervision does not need any manual annotation but it is learned on raw textual data. As our approach easily facilitates multilinguality, we evaluate it on datasets covering German (PHOENIX-2014T) and American (How2Sign) sign languages and experiment with mono- and multilingual sentence embeddings and translation systems. Our approach significantly outperforms other gloss-free approaches, setting the new state-of-the-art for data sets where glosses are not available and when no additional SLT datasets are used for pretraining, diminishing the gap between gloss-free and gloss-dependent systems.
title Sign Language Translation with Sentence Embedding Supervision
topic Computation and Language
url https://arxiv.org/abs/2510.19367