SignLLM: Sign Language Production Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Fang, Sen, Chen, Chen, Wang, Lei, Zheng, Ce, Sui, Chunyu, Tian, Yapeng
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915266075557888
author Fang, Sen
Chen, Chen
Wang, Lei
Zheng, Ce
Sui, Chunyu
Tian, Yapeng
author_facet Fang, Sen
Chen, Chen
Wang, Lei
Zheng, Ce
Sui, Chunyu
Tian, Yapeng
contents In this paper, we propose SignLLM, a multilingual Sign Language Production (SLP) large language model, which includes two novel multilingual SLP modes MLSF and Prompt2LangGloss that allow sign language gestures generation from query texts input and question-style prompts input respectively. Both modes can use a new RL loss based on reinforcement learning and a new RL module named Priority Learning Channel. These RL components can accelerate the training by enhancing the model's capability to sample high-quality data. To train SignLLM, we introduce Prompt2Sign, a comprehensive multilingual sign language dataset, which builds from public data, including American Sign Language (ASL) and seven others. This dataset standardizes information by extracting pose information from sign language videos into a unified compressed format. We extensively evaluate SignLLM, demonstrating that our model achieves state-of-the-art performance on SLP tasks across eight sign languages.
format Preprint
id arxiv_https___arxiv_org_abs_2405_10718
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SignLLM: Sign Language Production Large Language Models
Fang, Sen
Chen, Chen
Wang, Lei
Zheng, Ce
Sui, Chunyu
Tian, Yapeng
Computer Vision and Pattern Recognition
Computation and Language
In this paper, we propose SignLLM, a multilingual Sign Language Production (SLP) large language model, which includes two novel multilingual SLP modes MLSF and Prompt2LangGloss that allow sign language gestures generation from query texts input and question-style prompts input respectively. Both modes can use a new RL loss based on reinforcement learning and a new RL module named Priority Learning Channel. These RL components can accelerate the training by enhancing the model's capability to sample high-quality data. To train SignLLM, we introduce Prompt2Sign, a comprehensive multilingual sign language dataset, which builds from public data, including American Sign Language (ASL) and seven others. This dataset standardizes information by extracting pose information from sign language videos into a unified compressed format. We extensively evaluate SignLLM, demonstrating that our model achieves state-of-the-art performance on SLP tasks across eight sign languages.
title SignLLM: Sign Language Production Large Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2405.10718