Creating Personalized Synthetic Voices from Articulation Impaired Speech Using Augmented Reconstruction Loss

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tian, Yusheng, Li, Jingyu, Lee, Tan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929202482118656
author Tian, Yusheng
Li, Jingyu
Lee, Tan
author_facet Tian, Yusheng
Li, Jingyu
Lee, Tan
contents This research is about the creation of personalized synthetic voices for head and neck cancer survivors. It is focused particularly on tongue cancer patients whose speech might exhibit severe articulation impairment. Our goal is to restore normal articulation in the synthesized speech, while maximally preserving the target speaker's individuality in terms of both the voice timbre and speaking style. This is formulated as a task of learning from noisy labels. We propose to augment the commonly used speech reconstruction loss with two additional terms. The first term constitutes a regularization loss that mitigates the impact of distorted articulation in the training speech. The second term is a consistency loss that encourages correct articulation in the generated speech. These additional loss terms are obtained from frame-level articulation scores of original and generated speech, which are derived using a separately trained phone classifier. Experimental results on a real case of tongue cancer patient confirm that the synthetic voice achieves comparable articulation quality to unimpaired natural speech, while effectively maintaining the target speaker's individuality. Audio samples are available at https://myspeechproject.github.io/ArticulationRepair/.
format Preprint
id arxiv_https___arxiv_org_abs_2401_03816
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Creating Personalized Synthetic Voices from Articulation Impaired Speech Using Augmented Reconstruction Loss
Tian, Yusheng
Li, Jingyu
Lee, Tan
Audio and Speech Processing
Sound
This research is about the creation of personalized synthetic voices for head and neck cancer survivors. It is focused particularly on tongue cancer patients whose speech might exhibit severe articulation impairment. Our goal is to restore normal articulation in the synthesized speech, while maximally preserving the target speaker's individuality in terms of both the voice timbre and speaking style. This is formulated as a task of learning from noisy labels. We propose to augment the commonly used speech reconstruction loss with two additional terms. The first term constitutes a regularization loss that mitigates the impact of distorted articulation in the training speech. The second term is a consistency loss that encourages correct articulation in the generated speech. These additional loss terms are obtained from frame-level articulation scores of original and generated speech, which are derived using a separately trained phone classifier. Experimental results on a real case of tongue cancer patient confirm that the synthetic voice achieves comparable articulation quality to unimpaired natural speech, while effectively maintaining the target speaker's individuality. Audio samples are available at https://myspeechproject.github.io/ArticulationRepair/.
title Creating Personalized Synthetic Voices from Articulation Impaired Speech Using Augmented Reconstruction Loss
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2401.03816