Activation Steering for Accent-Neutralized Zero-Shot Text-To-Speech

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yang, Mu, Hansen, John H. L.
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914374798540800
author Yang, Mu
Hansen, John H. L.
author_facet Yang, Mu
Hansen, John H. L.
contents Zero-shot Text-to-Speech (TTS) models can generate speech that captures both the voice timbre and accent of a reference speaker. However, disentangling these attributes remains challenging, as the output often inherits both the accent and timbre from the reference. In this study, we introduce a novel, post-hoc, and training-free approach to neutralize accent while preserving the speaker's original timbre, utilizing inference-time activation steering. We first extract layer-specific "steering vectors" offline, which are derived from the internal activation differences within the TTS model between accented and native speech. During inference, the steering vectors are applied to guide the model to produce accent-neutralized, timbre-preserving speech. Empirical results demonstrate that the proposed steering vectors effectively mitigate the output accent and exhibit strong generalizability to unseen accented speakers, offering a practical solution for accent-free voice cloning.
format Preprint
id arxiv_https___arxiv_org_abs_2603_05977
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Activation Steering for Accent-Neutralized Zero-Shot Text-To-Speech
Yang, Mu
Hansen, John H. L.
Audio and Speech Processing
Zero-shot Text-to-Speech (TTS) models can generate speech that captures both the voice timbre and accent of a reference speaker. However, disentangling these attributes remains challenging, as the output often inherits both the accent and timbre from the reference. In this study, we introduce a novel, post-hoc, and training-free approach to neutralize accent while preserving the speaker's original timbre, utilizing inference-time activation steering. We first extract layer-specific "steering vectors" offline, which are derived from the internal activation differences within the TTS model between accented and native speech. During inference, the steering vectors are applied to guide the model to produce accent-neutralized, timbre-preserving speech. Empirical results demonstrate that the proposed steering vectors effectively mitigate the output accent and exhibit strong generalizability to unseen accented speakers, offering a practical solution for accent-free voice cloning.
title Activation Steering for Accent-Neutralized Zero-Shot Text-To-Speech
topic Audio and Speech Processing
url https://arxiv.org/abs/2603.05977