Quantifying Speaker Embedding Phonological Rule Interactions in Accented Speech Synthesis

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lertpetchpun, Thanathai, Lee, Yoonjeong, Trachu, Thanapat, Lee, Jihwan, Feng, Tiantian, Byrd, Dani, Narayanan, Shrikanth
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915758452244480
author Lertpetchpun, Thanathai
Lee, Yoonjeong
Trachu, Thanapat
Lee, Jihwan
Feng, Tiantian
Byrd, Dani
Narayanan, Shrikanth
author_facet Lertpetchpun, Thanathai
Lee, Yoonjeong
Trachu, Thanapat
Lee, Jihwan
Feng, Tiantian
Byrd, Dani
Narayanan, Shrikanth
contents Many spoken languages, including English, exhibit wide variation in dialects and accents, making accent control an important capability for flexible text-to-speech (TTS) models. Current TTS systems typically generate accented speech by conditioning on speaker embeddings associated with specific accents. While effective, this approach offers limited interpretability and controllability, as embeddings also encode traits such as timbre and emotion. In this study, we analyze the interaction between speaker embeddings and linguistically motivated phonological rules in accented speech synthesis. Using American and British English as a case study, we implement rules for flapping, rhoticity, and vowel correspondences. We propose the phoneme shift rate (PSR), a novel metric quantifying how strongly embeddings preserve or override rule-based transformations. Experiments show that combining rules with embeddings yields more authentic accents, while embeddings can attenuate or overwrite rules, revealing entanglement between accent and speaker identity. Our findings highlight rules as a lever for accent control and a framework for evaluating disentanglement in speech generation.
format Preprint
id arxiv_https___arxiv_org_abs_2601_14417
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Quantifying Speaker Embedding Phonological Rule Interactions in Accented Speech Synthesis
Lertpetchpun, Thanathai
Lee, Yoonjeong
Trachu, Thanapat
Lee, Jihwan
Feng, Tiantian
Byrd, Dani
Narayanan, Shrikanth
Computation and Language
Many spoken languages, including English, exhibit wide variation in dialects and accents, making accent control an important capability for flexible text-to-speech (TTS) models. Current TTS systems typically generate accented speech by conditioning on speaker embeddings associated with specific accents. While effective, this approach offers limited interpretability and controllability, as embeddings also encode traits such as timbre and emotion. In this study, we analyze the interaction between speaker embeddings and linguistically motivated phonological rules in accented speech synthesis. Using American and British English as a case study, we implement rules for flapping, rhoticity, and vowel correspondences. We propose the phoneme shift rate (PSR), a novel metric quantifying how strongly embeddings preserve or override rule-based transformations. Experiments show that combining rules with embeddings yields more authentic accents, while embeddings can attenuate or overwrite rules, revealing entanglement between accent and speaker identity. Our findings highlight rules as a lever for accent control and a framework for evaluating disentanglement in speech generation.
title Quantifying Speaker Embedding Phonological Rule Interactions in Accented Speech Synthesis
topic Computation and Language
url https://arxiv.org/abs/2601.14417