Large Sign Language Models: Toward 3D American Sign Language Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Sen, He, Xiaoxiao, Liu, Di, Xia, Zhaoyang, Zhao, Mingyu, Tan, Chaowei, Li, Vivian, Liu, Bo, Metaxas, Dimitris N., Kapadia, Mubbasir
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914151444512768
author Zhang, Sen
He, Xiaoxiao
Liu, Di
Xia, Zhaoyang
Zhao, Mingyu
Tan, Chaowei
Li, Vivian
Liu, Bo
Metaxas, Dimitris N.
Kapadia, Mubbasir
author_facet Zhang, Sen
He, Xiaoxiao
Liu, Di
Xia, Zhaoyang
Zhao, Mingyu
Tan, Chaowei
Li, Vivian
Liu, Bo
Metaxas, Dimitris N.
Kapadia, Mubbasir
contents We present Large Sign Language Models (LSLM), a novel framework for translating 3D American Sign Language (ASL) by leveraging Large Language Models (LLMs) as the backbone, which can benefit hearing-impaired individuals' virtual communication. Unlike existing sign language recognition methods that rely on 2D video, our approach directly utilizes 3D sign language data to capture rich spatial, gestural, and depth information in 3D scenes. This enables more accurate and resilient translation, enhancing digital communication accessibility for the hearing-impaired community. Beyond the task of ASL translation, our work explores the integration of complex, embodied multimodal languages into the processing capabilities of LLMs, moving beyond purely text-based inputs to broaden their understanding of human communication. We investigate both direct translation from 3D gesture features to text and an instruction-guided setting where translations can be modulated by external prompts, offering greater flexibility. This work provides a foundational step toward inclusive, multimodal intelligent systems capable of understanding diverse forms of language.
format Preprint
id arxiv_https___arxiv_org_abs_2511_08535
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Large Sign Language Models: Toward 3D American Sign Language Translation
Zhang, Sen
He, Xiaoxiao
Liu, Di
Xia, Zhaoyang
Zhao, Mingyu
Tan, Chaowei
Li, Vivian
Liu, Bo
Metaxas, Dimitris N.
Kapadia, Mubbasir
Computer Vision and Pattern Recognition
Artificial Intelligence
We present Large Sign Language Models (LSLM), a novel framework for translating 3D American Sign Language (ASL) by leveraging Large Language Models (LLMs) as the backbone, which can benefit hearing-impaired individuals' virtual communication. Unlike existing sign language recognition methods that rely on 2D video, our approach directly utilizes 3D sign language data to capture rich spatial, gestural, and depth information in 3D scenes. This enables more accurate and resilient translation, enhancing digital communication accessibility for the hearing-impaired community. Beyond the task of ASL translation, our work explores the integration of complex, embodied multimodal languages into the processing capabilities of LLMs, moving beyond purely text-based inputs to broaden their understanding of human communication. We investigate both direct translation from 3D gesture features to text and an instruction-guided setting where translations can be modulated by external prompts, offering greater flexibility. This work provides a foundational step toward inclusive, multimodal intelligent systems capable of understanding diverse forms of language.
title Large Sign Language Models: Toward 3D American Sign Language Translation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.08535