Efficient Emotion and Speaker Adaptation in LLM-Based TTS via Characteristic-Specific Partial Fine-Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Tianrui, Ge, Meng, Gong, Cheng, Qiang, Chunyu, Wang, Haoyu, Huang, Zikang, Jiang, Yu, Ni, Ye, Lu, Yuheng, Wang, Xiaobao, Chng, Engsiong, Chen, Xie, Wang, Longbiao, Dang, Jianwu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918374510952448
author Wang, Tianrui
Ge, Meng
Gong, Cheng
Qiang, Chunyu
Wang, Haoyu
Huang, Zikang
Jiang, Yu
Ni, Ye
Lu, Yuheng
Wang, Xiaobao
Chng, Engsiong
Chen, Xie
Wang, Longbiao
Dang, Jianwu
author_facet Wang, Tianrui
Ge, Meng
Gong, Cheng
Qiang, Chunyu
Wang, Haoyu
Huang, Zikang
Jiang, Yu
Ni, Ye
Lu, Yuheng
Wang, Xiaobao
Chng, Engsiong
Chen, Xie
Wang, Longbiao
Dang, Jianwu
contents While LLM-based TTS models exhibit zero-shot emotion and speaker cloning, their cloning fidelity and pronunciation clarity degrade on unseen domains. Fine-tuning is essential for adaptation, yet uniform approaches overlook specific parameter contributions. Uniform tuning on limited data causes slow training and catastrophic forgetting, leading to degraded pronunciation accuracy. To address this, we propose CSP-FT, a characteristic-specific partial fine-tuning strategy. By dynamically analyzing layer contributions via a weighted sum, we selectively fine-tune only the two layers capturing the most and least emotion and speaker information, maximizing the utility of the former while explicitly strengthening the capacity of the latter. Experiments on a combined corpus of 11 datasets show CSP-FT matches or exceeds the fidelity and intelligibility of full fine-tuning while updating only ~8% of parameters, accelerating training by ~2x, and significantly mitigating catastrophic forgetting.
format Preprint
id arxiv_https___arxiv_org_abs_2501_14273
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Emotion and Speaker Adaptation in LLM-Based TTS via Characteristic-Specific Partial Fine-Tuning
Wang, Tianrui
Ge, Meng
Gong, Cheng
Qiang, Chunyu
Wang, Haoyu
Huang, Zikang
Jiang, Yu
Ni, Ye
Lu, Yuheng
Wang, Xiaobao
Chng, Engsiong
Chen, Xie
Wang, Longbiao
Dang, Jianwu
Audio and Speech Processing
Sound
While LLM-based TTS models exhibit zero-shot emotion and speaker cloning, their cloning fidelity and pronunciation clarity degrade on unseen domains. Fine-tuning is essential for adaptation, yet uniform approaches overlook specific parameter contributions. Uniform tuning on limited data causes slow training and catastrophic forgetting, leading to degraded pronunciation accuracy. To address this, we propose CSP-FT, a characteristic-specific partial fine-tuning strategy. By dynamically analyzing layer contributions via a weighted sum, we selectively fine-tune only the two layers capturing the most and least emotion and speaker information, maximizing the utility of the former while explicitly strengthening the capacity of the latter. Experiments on a combined corpus of 11 datasets show CSP-FT matches or exceeds the fidelity and intelligibility of full fine-tuning while updating only ~8% of parameters, accelerating training by ~2x, and significantly mitigating catastrophic forgetting.
title Efficient Emotion and Speaker Adaptation in LLM-Based TTS via Characteristic-Specific Partial Fine-Tuning
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2501.14273