Acoustic and perceptual differences between standard and accented speech and their voice clones

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yang, Tianle, Sun, Chengzhe, Rose, Phil, Lyu, Siwei
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918534212222976
author Yang, Tianle
Sun, Chengzhe
Rose, Phil
Lyu, Siwei
author_facet Yang, Tianle
Sun, Chengzhe
Rose, Phil
Lyu, Siwei
contents Voice cloning is often evaluated in terms of overall quality, but less is known about accent preservation and its perceptual consequences. We compare standard and heavily accented Mandarin speech and their voice clones using a combined computational and perceptual design. Embedding-based analyses showed larger original-clone distances for accented speakers in several speaker-discriminative embedding spaces, but this difference disappeared after normalizing against each speaker's within-original baseline variability. In the perception study, clones are rated as more similar to their originals for standard than for accented speakers, and intelligibility increases from original to clone, with a larger gain for accented speech. These results show that accent variation can shape perceived identity match and intelligibility in voice cloning even when it is not reflected in baseline-normalized speaker-embedding distance, and they motivate treating accent preservation as an explicit component of speaker identity preservation, rather than assuming that it is fully captured by off-the-shelf speaker-discriminative embeddings.
format Preprint
id arxiv_https___arxiv_org_abs_2604_01562
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Acoustic and perceptual differences between standard and accented speech and their voice clones
Yang, Tianle
Sun, Chengzhe
Rose, Phil
Lyu, Siwei
Sound
Artificial Intelligence
Computation and Language
Computers and Society
Human-Computer Interaction
Voice cloning is often evaluated in terms of overall quality, but less is known about accent preservation and its perceptual consequences. We compare standard and heavily accented Mandarin speech and their voice clones using a combined computational and perceptual design. Embedding-based analyses showed larger original-clone distances for accented speakers in several speaker-discriminative embedding spaces, but this difference disappeared after normalizing against each speaker's within-original baseline variability. In the perception study, clones are rated as more similar to their originals for standard than for accented speakers, and intelligibility increases from original to clone, with a larger gain for accented speech. These results show that accent variation can shape perceived identity match and intelligibility in voice cloning even when it is not reflected in baseline-normalized speaker-embedding distance, and they motivate treating accent preservation as an explicit component of speaker identity preservation, rather than assuming that it is fully captured by off-the-shelf speaker-discriminative embeddings.
title Acoustic and perceptual differences between standard and accented speech and their voice clones
topic Sound
Artificial Intelligence
Computation and Language
Computers and Society
Human-Computer Interaction
url https://arxiv.org/abs/2604.01562