Score-Based Training for Energy-Based TTS Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Wanli, Ragni, Anton
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909616954146816
author Sun, Wanli
Ragni, Anton
author_facet Sun, Wanli
Ragni, Anton
contents Noise contrastive estimation (NCE) is a popular method for training energy-based models (EBM) with intractable normalisation terms. The key idea of NCE is to learn by comparing unnormalised log-likelihoods of the reference and noisy samples, thus avoiding explicitly computing normalisation terms. However, NCE critically relies on the quality of noisy samples. Recently, sliced score matching (SSM) has been popularised by closely related diffusion models (DM). Unlike NCE, SSM learns a gradient of log-likelihood, or score, by learning distribution of its projections on randomly chosen directions. However, both NCE and SSM disregard the form of log-likelihood function, which is problematic given that EBMs and DMs make use of first-order optimisation during inference. This paper proposes a new criterion that learns scores more suitable for first-order schemes. Experiments contrasts these approaches for training EBMs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13771
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Score-Based Training for Energy-Based TTS Models
Sun, Wanli
Ragni, Anton
Sound
Machine Learning
Audio and Speech Processing
Noise contrastive estimation (NCE) is a popular method for training energy-based models (EBM) with intractable normalisation terms. The key idea of NCE is to learn by comparing unnormalised log-likelihoods of the reference and noisy samples, thus avoiding explicitly computing normalisation terms. However, NCE critically relies on the quality of noisy samples. Recently, sliced score matching (SSM) has been popularised by closely related diffusion models (DM). Unlike NCE, SSM learns a gradient of log-likelihood, or score, by learning distribution of its projections on randomly chosen directions. However, both NCE and SSM disregard the form of log-likelihood function, which is problematic given that EBMs and DMs make use of first-order optimisation during inference. This paper proposes a new criterion that learns scores more suitable for first-order schemes. Experiments contrasts these approaches for training EBMs.
title Score-Based Training for Energy-Based TTS Models
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2505.13771