SignAttention: On the Interpretability of Transformer Models for Sign Language Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bianco, Pedro Alejandro Dal, Stanchi, Oscar Agustín, Quiroga, Facundo Manuel, Ronchetti, Franco, Ferrante, Enzo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913553612537856
author Bianco, Pedro Alejandro Dal
Stanchi, Oscar Agustín
Quiroga, Facundo Manuel
Ronchetti, Franco
Ferrante, Enzo
author_facet Bianco, Pedro Alejandro Dal
Stanchi, Oscar Agustín
Quiroga, Facundo Manuel
Ronchetti, Franco
Ferrante, Enzo
contents This paper presents the first comprehensive interpretability analysis of a Transformer-based Sign Language Translation (SLT) model, focusing on the translation from video-based Greek Sign Language to glosses and text. Leveraging the Greek Sign Language Dataset, we examine the attention mechanisms within the model to understand how it processes and aligns visual input with sequential glosses. Our analysis reveals that the model pays attention to clusters of frames rather than individual ones, with a diagonal alignment pattern emerging between poses and glosses, which becomes less distinct as the number of glosses increases. We also explore the relative contributions of cross-attention and self-attention at each decoding step, finding that the model initially relies on video frames but shifts its focus to previously predicted tokens as the translation progresses. This work contributes to a deeper understanding of SLT models, paving the way for the development of more transparent and reliable translation systems essential for real-world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14506
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SignAttention: On the Interpretability of Transformer Models for Sign Language Translation
Bianco, Pedro Alejandro Dal
Stanchi, Oscar Agustín
Quiroga, Facundo Manuel
Ronchetti, Franco
Ferrante, Enzo
Computation and Language
Artificial Intelligence
This paper presents the first comprehensive interpretability analysis of a Transformer-based Sign Language Translation (SLT) model, focusing on the translation from video-based Greek Sign Language to glosses and text. Leveraging the Greek Sign Language Dataset, we examine the attention mechanisms within the model to understand how it processes and aligns visual input with sequential glosses. Our analysis reveals that the model pays attention to clusters of frames rather than individual ones, with a diagonal alignment pattern emerging between poses and glosses, which becomes less distinct as the number of glosses increases. We also explore the relative contributions of cross-attention and self-attention at each decoding step, finding that the model initially relies on video frames but shifts its focus to previously predicted tokens as the translation progresses. This work contributes to a deeper understanding of SLT models, paving the way for the development of more transparent and reliable translation systems essential for real-world applications.
title SignAttention: On the Interpretability of Transformer Models for Sign Language Translation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.14506