Gloss-Free Sign Language Translation: An Unbiased Evaluation of Progress in the Field

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sincan, Ozge Mercanoglu, Low, Jian He, Asasi, Sobhan, Bowden, Richard
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914392406228992
author Sincan, Ozge Mercanoglu
Low, Jian He
Asasi, Sobhan
Bowden, Richard
author_facet Sincan, Ozge Mercanoglu
Low, Jian He
Asasi, Sobhan
Bowden, Richard
contents Sign Language Translation (SLT) aims to automatically convert visual sign language videos into spoken language text and vice versa. While recent years have seen rapid progress, the true sources of performance improvements often remain unclear. Do reported performance gains come from methodological novelty, or from the choice of a different backbone, training optimizations, hyperparameter tuning, or even differences in the calculation of evaluation metrics? This paper presents a comprehensive study of recent gloss-free SLT models by re-implementing key contributions in a unified codebase. We ensure fair comparison by standardizing preprocessing, video encoders, and training setups across all methods. Our analysis shows that many of the performance gains reported in the literature often diminish when models are evaluated under consistent conditions, suggesting that implementation details and evaluation setups play a significant role in determining results. We make the codebase publicly available here (https://github.com/ozgemercanoglu/sltbaselines) to support transparency and reproducibility in SLT research.
format Preprint
id arxiv_https___arxiv_org_abs_2603_13240
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Gloss-Free Sign Language Translation: An Unbiased Evaluation of Progress in the Field
Sincan, Ozge Mercanoglu
Low, Jian He
Asasi, Sobhan
Bowden, Richard
Computer Vision and Pattern Recognition
Computation and Language
Sign Language Translation (SLT) aims to automatically convert visual sign language videos into spoken language text and vice versa. While recent years have seen rapid progress, the true sources of performance improvements often remain unclear. Do reported performance gains come from methodological novelty, or from the choice of a different backbone, training optimizations, hyperparameter tuning, or even differences in the calculation of evaluation metrics? This paper presents a comprehensive study of recent gloss-free SLT models by re-implementing key contributions in a unified codebase. We ensure fair comparison by standardizing preprocessing, video encoders, and training setups across all methods. Our analysis shows that many of the performance gains reported in the literature often diminish when models are evaluated under consistent conditions, suggesting that implementation details and evaluation setups play a significant role in determining results. We make the codebase publicly available here (https://github.com/ozgemercanoglu/sltbaselines) to support transparency and reproducibility in SLT research.
title Gloss-Free Sign Language Translation: An Unbiased Evaluation of Progress in the Field
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2603.13240