MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xia, Kangxiang, Zhu, Xinfa, Yao, Jixun, Xie, Lei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912562546737152
author Xia, Kangxiang
Zhu, Xinfa
Yao, Jixun
Xie, Lei
author_facet Xia, Kangxiang
Zhu, Xinfa
Yao, Jixun
Xie, Lei
contents In recent years, text-to-speech (TTS) has seen impressive advancements through large-scale language models, achieving human-level speech quality. Integrating human feedback has proven effective for enhancing robustness in these systems. However, current approaches face challenges in optimizing TTS with preference data across multiple dimensions and often suffer from performance degradation due to overconfidence in rewards. We propose Multidimensional Preference Optimization (MPO) to better align TTS systems with human preferences. MPO introduces a preference set that streamlines the construction of data for multidimensional preference optimization, enabling alignment with multiple dimensions. Additionally, we incorporate regularization during training to address the typical degradation issues in DPO-based approaches. Our experiments demonstrate MPO's effectiveness, showing significant improvements in intelligibility, speaker similarity, and prosody compared to baseline systems.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00685
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
Xia, Kangxiang
Zhu, Xinfa
Yao, Jixun
Xie, Lei
Audio and Speech Processing
Sound
In recent years, text-to-speech (TTS) has seen impressive advancements through large-scale language models, achieving human-level speech quality. Integrating human feedback has proven effective for enhancing robustness in these systems. However, current approaches face challenges in optimizing TTS with preference data across multiple dimensions and often suffer from performance degradation due to overconfidence in rewards. We propose Multidimensional Preference Optimization (MPO) to better align TTS systems with human preferences. MPO introduces a preference set that streamlines the construction of data for multidimensional preference optimization, enabling alignment with multiple dimensions. Additionally, we incorporate regularization during training to address the typical degradation issues in DPO-based approaches. Our experiments demonstrate MPO's effectiveness, showing significant improvements in intelligibility, speaker similarity, and prosody compared to baseline systems.
title MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2509.00685