Cued Speech Generation Leveraging a Pre-trained Audiovisual Text-to-Speech Model

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Sankar, Sanjana, Lenglet, Martin, Bailly, Gerard, Beautemps, Denis, Hueber, Thomas
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915095577100288
author Sankar, Sanjana
Lenglet, Martin
Bailly, Gerard
Beautemps, Denis
Hueber, Thomas
author_facet Sankar, Sanjana
Lenglet, Martin
Bailly, Gerard
Beautemps, Denis
Hueber, Thomas
contents This paper presents a novel approach for the automatic generation of Cued Speech (ACSG), a visual communication system used by people with hearing impairment to better elicit the spoken language. We explore transfer learning strategies by leveraging a pre-trained audiovisual autoregressive text-to-speech model (AVTacotron2). This model is reprogrammed to infer Cued Speech (CS) hand and lip movements from text input. Experiments are conducted on two publicly available datasets, including one recorded specifically for this study. Performance is assessed using an automatic CS recognition system. With a decoding accuracy at the phonetic level reaching approximately 77%, the results demonstrate the effectiveness of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2501_04799
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Cued Speech Generation Leveraging a Pre-trained Audiovisual Text-to-Speech Model
Sankar, Sanjana
Lenglet, Martin
Bailly, Gerard
Beautemps, Denis
Hueber, Thomas
Computation and Language
This paper presents a novel approach for the automatic generation of Cued Speech (ACSG), a visual communication system used by people with hearing impairment to better elicit the spoken language. We explore transfer learning strategies by leveraging a pre-trained audiovisual autoregressive text-to-speech model (AVTacotron2). This model is reprogrammed to infer Cued Speech (CS) hand and lip movements from text input. Experiments are conducted on two publicly available datasets, including one recorded specifically for this study. Performance is assessed using an automatic CS recognition system. With a decoding accuracy at the phonetic level reaching approximately 77%, the results demonstrate the effectiveness of our approach.
title Cued Speech Generation Leveraging a Pre-trained Audiovisual Text-to-Speech Model
topic Computation and Language
url https://arxiv.org/abs/2501.04799