Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Kyowoon, Stitsyuk, Artyom, Jho, Gunu, Hwang, Inchul, Choi, Jaesik
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916771816013824
author Lee, Kyowoon
Stitsyuk, Artyom
Jho, Gunu
Hwang, Inchul
Choi, Jaesik
author_facet Lee, Kyowoon
Stitsyuk, Artyom
Jho, Gunu
Hwang, Inchul
Choi, Jaesik
contents Recent advances in Text-to-Speech (TTS) have significantly improved speech naturalness, increasing the demand for precise prosody control and mispronunciation correction. Existing approaches for prosody manipulation often depend on specialized modules or additional training, limiting their capacity for post-hoc adjustments. Similarly, traditional mispronunciation correction relies on grapheme-to-phoneme dictionaries, making it less practical in low-resource settings. We introduce Counterfactual Activation Editing, a model-agnostic method that manipulates internal representations in a pre-trained TTS model to achieve post-hoc control of prosody and pronunciation. Experimental results show that our method effectively adjusts prosodic features and corrects mispronunciations while preserving synthesis quality. This opens the door to inference-time refinement of TTS outputs without retraining, bridging the gap between pre-trained TTS models and editable speech synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00832
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models
Lee, Kyowoon
Stitsyuk, Artyom
Jho, Gunu
Hwang, Inchul
Choi, Jaesik
Sound
Artificial Intelligence
Audio and Speech Processing
Recent advances in Text-to-Speech (TTS) have significantly improved speech naturalness, increasing the demand for precise prosody control and mispronunciation correction. Existing approaches for prosody manipulation often depend on specialized modules or additional training, limiting their capacity for post-hoc adjustments. Similarly, traditional mispronunciation correction relies on grapheme-to-phoneme dictionaries, making it less practical in low-resource settings. We introduce Counterfactual Activation Editing, a model-agnostic method that manipulates internal representations in a pre-trained TTS model to achieve post-hoc control of prosody and pronunciation. Experimental results show that our method effectively adjusts prosodic features and corrects mispronunciations while preserving synthesis quality. This opens the door to inference-time refinement of TTS outputs without retraining, bridging the gap between pre-trained TTS models and editable speech synthesis.
title Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2506.00832