MultimodalHugs: Enabling Sign Language Processing in Hugging Face

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sant, Gerard, Jiang, Zifan, Escolano, Carlos, Moryossef, Amit, Müller, Mathias, Sennrich, Rico, Ebling, Sarah
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915491310731264
author Sant, Gerard
Jiang, Zifan
Escolano, Carlos
Moryossef, Amit
Müller, Mathias
Sennrich, Rico
Ebling, Sarah
author_facet Sant, Gerard
Jiang, Zifan
Escolano, Carlos
Moryossef, Amit
Müller, Mathias
Sennrich, Rico
Ebling, Sarah
contents In recent years, sign language processing (SLP) has gained importance in the general field of Natural Language Processing. However, compared to research on spoken languages, SLP research is hindered by complex ad-hoc code, inadvertently leading to low reproducibility and unfair comparisons. Existing tools that are built for fast and reproducible experimentation, such as Hugging Face, are not flexible enough to seamlessly integrate sign language experiments. This view is confirmed by a survey we conducted among SLP researchers. To address these challenges, we introduce MultimodalHugs, a framework built on top of Hugging Face that enables more diverse data modalities and tasks, while inheriting the well-known advantages of the Hugging Face ecosystem. Even though sign languages are our primary focus, MultimodalHugs adds a layer of abstraction that makes it more widely applicable to other use cases that do not fit one of the standard templates of Hugging Face. We provide quantitative experiments to illustrate how MultimodalHugs can accommodate diverse modalities such as pose estimation data for sign languages, or pixel data for text characters.
format Preprint
id arxiv_https___arxiv_org_abs_2509_09729
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MultimodalHugs: Enabling Sign Language Processing in Hugging Face
Sant, Gerard
Jiang, Zifan
Escolano, Carlos
Moryossef, Amit
Müller, Mathias
Sennrich, Rico
Ebling, Sarah
Computation and Language
Artificial Intelligence
Multimedia
In recent years, sign language processing (SLP) has gained importance in the general field of Natural Language Processing. However, compared to research on spoken languages, SLP research is hindered by complex ad-hoc code, inadvertently leading to low reproducibility and unfair comparisons. Existing tools that are built for fast and reproducible experimentation, such as Hugging Face, are not flexible enough to seamlessly integrate sign language experiments. This view is confirmed by a survey we conducted among SLP researchers. To address these challenges, we introduce MultimodalHugs, a framework built on top of Hugging Face that enables more diverse data modalities and tasks, while inheriting the well-known advantages of the Hugging Face ecosystem. Even though sign languages are our primary focus, MultimodalHugs adds a layer of abstraction that makes it more widely applicable to other use cases that do not fit one of the standard templates of Hugging Face. We provide quantitative experiments to illustrate how MultimodalHugs can accommodate diverse modalities such as pose estimation data for sign languages, or pixel data for text characters.
title MultimodalHugs: Enabling Sign Language Processing in Hugging Face
topic Computation and Language
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2509.09729