Exploring Pose-based Sign Language Translation: Ablation Studies and Attention Insights

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zelezny, Tomas, Straka, Jakub, Javorek, Vaclav, Valach, Ondrej, Hruz, Marek, Gruber, Ivan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916821524807680
author Zelezny, Tomas
Straka, Jakub
Javorek, Vaclav
Valach, Ondrej
Hruz, Marek
Gruber, Ivan
author_facet Zelezny, Tomas
Straka, Jakub
Javorek, Vaclav
Valach, Ondrej
Hruz, Marek
Gruber, Ivan
contents Sign Language Translation (SLT) has evolved significantly, moving from isolated recognition approaches to complex, continuous gloss-free translation systems. This paper explores the impact of pose-based data preprocessing techniques - normalization, interpolation, and augmentation - on SLT performance. We employ a transformer-based architecture, adapting a modified T5 encoder-decoder model to process pose representations. Through extensive ablation studies on YouTubeASL and How2Sign datasets, we analyze how different preprocessing strategies affect translation accuracy. Our results demonstrate that appropriate normalization, interpolation, and augmentation techniques can significantly improve model robustness and generalization abilities. Additionally, we provide a deep analysis of the model's attentions and reveal interesting behavior suggesting that adding a dedicated register token can improve overall model performance. We publish our code on our GitHub repository, including the preprocessed YouTubeASL data.
format Preprint
id arxiv_https___arxiv_org_abs_2507_01532
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring Pose-based Sign Language Translation: Ablation Studies and Attention Insights
Zelezny, Tomas
Straka, Jakub
Javorek, Vaclav
Valach, Ondrej
Hruz, Marek
Gruber, Ivan
Computer Vision and Pattern Recognition
Sign Language Translation (SLT) has evolved significantly, moving from isolated recognition approaches to complex, continuous gloss-free translation systems. This paper explores the impact of pose-based data preprocessing techniques - normalization, interpolation, and augmentation - on SLT performance. We employ a transformer-based architecture, adapting a modified T5 encoder-decoder model to process pose representations. Through extensive ablation studies on YouTubeASL and How2Sign datasets, we analyze how different preprocessing strategies affect translation accuracy. Our results demonstrate that appropriate normalization, interpolation, and augmentation techniques can significantly improve model robustness and generalization abilities. Additionally, we provide a deep analysis of the model's attentions and reveal interesting behavior suggesting that adding a dedicated register token can improve overall model performance. We publish our code on our GitHub repository, including the preprocessed YouTubeASL data.
title Exploring Pose-based Sign Language Translation: Ablation Studies and Attention Insights
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.01532