Towards Inclusive Communication: A Unified Framework for Generating Spoken Language from Sign, Lip, and Audio

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yeo, Jeong Hun, Rha, Hyeongseop, Park, Sungjune, Won, Junil, Ro, Yong Man
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914415457075200
author Yeo, Jeong Hun
Rha, Hyeongseop
Park, Sungjune
Won, Junil
Ro, Yong Man
author_facet Yeo, Jeong Hun
Rha, Hyeongseop
Park, Sungjune
Won, Junil
Ro, Yong Man
contents Audio is the primary modality for human communication and has driven the success of Automatic Speech Recognition (ASR) technologies. However, such audio-centric systems inherently exclude individuals who are deaf or hard of hearing. Visual alternatives such as sign language and lip reading offer effective substitutes, and recent advances in Sign Language Translation (SLT) and Visual Speech Recognition (VSR) have improved audio-less communication. Yet, these modalities have largely been studied in isolation, and their integration within a unified framework remains underexplored. In this paper, we propose the first unified framework capable of handling diverse combinations of sign language, lip movements, and audio for spoken-language text generation. We focus on three main objectives: (i) designing a unified, modality-agnostic architecture capable of effectively processing heterogeneous inputs; (ii) exploring the underexamined synergy among modalities, particularly the role of lip movements as non-manual cues in sign language comprehension; and (iii) achieving performance on par with or superior to state-of-the-art models specialized for individual tasks. Building on this framework, we achieve performance on par with or better than task-specific state-of-the-art models across SLT, VSR, ASR, and Audio-Visual Speech Recognition. Furthermore, our analysis reveals a key linguistic insight: explicitly modeling lip movements as a distinct modality significantly improves SLT performance by capturing critical non-manual cues.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20476
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Inclusive Communication: A Unified Framework for Generating Spoken Language from Sign, Lip, and Audio
Yeo, Jeong Hun
Rha, Hyeongseop
Park, Sungjune
Won, Junil
Ro, Yong Man
Computer Vision and Pattern Recognition
Multimedia
Audio and Speech Processing
Image and Video Processing
Audio is the primary modality for human communication and has driven the success of Automatic Speech Recognition (ASR) technologies. However, such audio-centric systems inherently exclude individuals who are deaf or hard of hearing. Visual alternatives such as sign language and lip reading offer effective substitutes, and recent advances in Sign Language Translation (SLT) and Visual Speech Recognition (VSR) have improved audio-less communication. Yet, these modalities have largely been studied in isolation, and their integration within a unified framework remains underexplored. In this paper, we propose the first unified framework capable of handling diverse combinations of sign language, lip movements, and audio for spoken-language text generation. We focus on three main objectives: (i) designing a unified, modality-agnostic architecture capable of effectively processing heterogeneous inputs; (ii) exploring the underexamined synergy among modalities, particularly the role of lip movements as non-manual cues in sign language comprehension; and (iii) achieving performance on par with or superior to state-of-the-art models specialized for individual tasks. Building on this framework, we achieve performance on par with or better than task-specific state-of-the-art models across SLT, VSR, ASR, and Audio-Visual Speech Recognition. Furthermore, our analysis reveals a key linguistic insight: explicitly modeling lip movements as a distinct modality significantly improves SLT performance by capturing critical non-manual cues.
title Towards Inclusive Communication: A Unified Framework for Generating Spoken Language from Sign, Lip, and Audio
topic Computer Vision and Pattern Recognition
Multimedia
Audio and Speech Processing
Image and Video Processing
url https://arxiv.org/abs/2508.20476