Leveraging the Power of MLLMs for Gloss-Free Sign Language Translation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kim, Jungeun, Jeon, Hyeongwoo, Bae, Jongseong, Kim, Ha Young
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909751965646848
author Kim, Jungeun
Jeon, Hyeongwoo
Bae, Jongseong
Kim, Ha Young
author_facet Kim, Jungeun
Jeon, Hyeongwoo
Bae, Jongseong
Kim, Ha Young
contents Sign language translation (SLT) is a challenging task that involves translating sign language images into spoken language. For SLT models to perform this task successfully, they must bridge the modality gap and identify subtle variations in sign language components to understand their meanings accurately. To address these challenges, we propose a novel gloss-free SLT framework called Multimodal Sign Language Translation (MMSLT), which leverages the representational capabilities of off-the-shelf multimodal large language models (MLLMs). Specifically, we use MLLMs to generate detailed textual descriptions of sign language components. Then, through our proposed multimodal-language pre-training module, we integrate these description features with sign video features to align them within the spoken sentence space. Our approach achieves state-of-the-art performance on benchmark datasets PHOENIX14T and CSL-Daily, highlighting the potential of MLLMs to be utilized effectively in SLT. Code is available at https://github.com/hwjeon98/MMSLT.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16789
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Leveraging the Power of MLLMs for Gloss-Free Sign Language Translation
Kim, Jungeun
Jeon, Hyeongwoo
Bae, Jongseong
Kim, Ha Young
Computer Vision and Pattern Recognition
Computation and Language
Sign language translation (SLT) is a challenging task that involves translating sign language images into spoken language. For SLT models to perform this task successfully, they must bridge the modality gap and identify subtle variations in sign language components to understand their meanings accurately. To address these challenges, we propose a novel gloss-free SLT framework called Multimodal Sign Language Translation (MMSLT), which leverages the representational capabilities of off-the-shelf multimodal large language models (MLLMs). Specifically, we use MLLMs to generate detailed textual descriptions of sign language components. Then, through our proposed multimodal-language pre-training module, we integrate these description features with sign video features to align them within the spoken sentence space. Our approach achieves state-of-the-art performance on benchmark datasets PHOENIX14T and CSL-Daily, highlighting the potential of MLLMs to be utilized effectively in SLT. Code is available at https://github.com/hwjeon98/MMSLT.
title Leveraging the Power of MLLMs for Gloss-Free Sign Language Translation
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2411.16789