VAE-based Phoneme Alignment Using Gradient Annealing and SSL Acoustic Features

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Koriyama, Tomoki
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910619000635392
author Koriyama, Tomoki
author_facet Koriyama, Tomoki
contents This paper presents an accurate phoneme alignment model that aims for speech analysis and video content creation. We propose a variational autoencoder (VAE)-based alignment model in which a probable path is searched using encoded acoustic and linguistic embeddings in an unsupervised manner. Our proposed model is based on one TTS alignment (OTA) and extended to obtain phoneme boundaries. Specifically, we incorporate a VAE architecture to maintain consistency between the embedding and input, apply gradient annealing to avoid local optimum during training, and introduce a self-supervised learning (SSL)-based acoustic-feature input and state-level linguistic unit to utilize rich and detailed information. Experimental results show that the proposed model generated phoneme boundaries closer to annotated ones compared with the conventional OTA model, the CTC-based segmentation model, and the widely-used tool MFA.
format Preprint
id arxiv_https___arxiv_org_abs_2407_02749
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VAE-based Phoneme Alignment Using Gradient Annealing and SSL Acoustic Features
Koriyama, Tomoki
Audio and Speech Processing
Sound
This paper presents an accurate phoneme alignment model that aims for speech analysis and video content creation. We propose a variational autoencoder (VAE)-based alignment model in which a probable path is searched using encoded acoustic and linguistic embeddings in an unsupervised manner. Our proposed model is based on one TTS alignment (OTA) and extended to obtain phoneme boundaries. Specifically, we incorporate a VAE architecture to maintain consistency between the embedding and input, apply gradient annealing to avoid local optimum during training, and introduce a self-supervised learning (SSL)-based acoustic-feature input and state-level linguistic unit to utilize rich and detailed information. Experimental results show that the proposed model generated phoneme boundaries closer to annotated ones compared with the conventional OTA model, the CTC-based segmentation model, and the widely-used tool MFA.
title VAE-based Phoneme Alignment Using Gradient Annealing and SSL Acoustic Features
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2407.02749