Eigenvoice Synthesis based on Model Editing for Speaker Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Murata, Masato, Miyazaki, Koichi, Koriyama, Tomoki, Toda, Tomoki
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918083249045504
author Murata, Masato
Miyazaki, Koichi
Koriyama, Tomoki
Toda, Tomoki
author_facet Murata, Masato
Miyazaki, Koichi
Koriyama, Tomoki
Toda, Tomoki
contents Speaker generation task aims to create unseen speaker voice without reference speech. The key to the task is defining a speaker space that represents diverse speakers to determine the generated speaker trait. However, the effective way to define this speaker space remains unclear. Eigenvoice synthesis is one of the promising approaches in the traditional parametric synthesis framework, such as HMM-based methods, which define a low-dimensional speaker space using pre-stored speaker features. This study proposes a novel DNN-based eigenvoice synthesis method via model editing. Unlike prior methods, our method defines a speaker space in the DNN model parameter space. By directly sampling new DNN model parameters in this space, we can create diverse speaker voices. Experimental results showed the capability of our method to generate diverse speakers' speech. Moreover, we discovered a gender-dominant axis in the created speaker space, indicating the potential to control speaker attributes.
format Preprint
id arxiv_https___arxiv_org_abs_2507_03377
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Eigenvoice Synthesis based on Model Editing for Speaker Generation
Murata, Masato
Miyazaki, Koichi
Koriyama, Tomoki
Toda, Tomoki
Sound
Audio and Speech Processing
Speaker generation task aims to create unseen speaker voice without reference speech. The key to the task is defining a speaker space that represents diverse speakers to determine the generated speaker trait. However, the effective way to define this speaker space remains unclear. Eigenvoice synthesis is one of the promising approaches in the traditional parametric synthesis framework, such as HMM-based methods, which define a low-dimensional speaker space using pre-stored speaker features. This study proposes a novel DNN-based eigenvoice synthesis method via model editing. Unlike prior methods, our method defines a speaker space in the DNN model parameter space. By directly sampling new DNN model parameters in this space, we can create diverse speaker voices. Experimental results showed the capability of our method to generate diverse speakers' speech. Moreover, we discovered a gender-dominant axis in the created speaker space, indicating the potential to control speaker attributes.
title Eigenvoice Synthesis based on Model Editing for Speaker Generation
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2507.03377