Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916771816013824 |
|---|---|
| author | Lee, Kyowoon Stitsyuk, Artyom Jho, Gunu Hwang, Inchul Choi, Jaesik |
| author_facet | Lee, Kyowoon Stitsyuk, Artyom Jho, Gunu Hwang, Inchul Choi, Jaesik |
| contents | Recent advances in Text-to-Speech (TTS) have significantly improved speech naturalness, increasing the demand for precise prosody control and mispronunciation correction. Existing approaches for prosody manipulation often depend on specialized modules or additional training, limiting their capacity for post-hoc adjustments. Similarly, traditional mispronunciation correction relies on grapheme-to-phoneme dictionaries, making it less practical in low-resource settings. We introduce Counterfactual Activation Editing, a model-agnostic method that manipulates internal representations in a pre-trained TTS model to achieve post-hoc control of prosody and pronunciation. Experimental results show that our method effectively adjusts prosodic features and corrects mispronunciations while preserving synthesis quality. This opens the door to inference-time refinement of TTS outputs without retraining, bridging the gap between pre-trained TTS models and editable speech synthesis. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_00832 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models Lee, Kyowoon Stitsyuk, Artyom Jho, Gunu Hwang, Inchul Choi, Jaesik Sound Artificial Intelligence Audio and Speech Processing Recent advances in Text-to-Speech (TTS) have significantly improved speech naturalness, increasing the demand for precise prosody control and mispronunciation correction. Existing approaches for prosody manipulation often depend on specialized modules or additional training, limiting their capacity for post-hoc adjustments. Similarly, traditional mispronunciation correction relies on grapheme-to-phoneme dictionaries, making it less practical in low-resource settings. We introduce Counterfactual Activation Editing, a model-agnostic method that manipulates internal representations in a pre-trained TTS model to achieve post-hoc control of prosody and pronunciation. Experimental results show that our method effectively adjusts prosodic features and corrects mispronunciations while preserving synthesis quality. This opens the door to inference-time refinement of TTS outputs without retraining, bridging the gap between pre-trained TTS models and editable speech synthesis. |
| title | Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models |
| topic | Sound Artificial Intelligence Audio and Speech Processing |
| url | https://arxiv.org/abs/2506.00832 |