The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Jian, Yang, Changbing, Samir, Farhan, Islam, Jahurul
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914738171019264
author Zhu, Jian
Yang, Changbing
Samir, Farhan
Islam, Jahurul
author_facet Zhu, Jian
Yang, Changbing
Samir, Farhan
Islam, Jahurul
contents In this project, we demonstrate that phoneme-based models for speech processing can achieve strong crosslinguistic generalizability to unseen languages. We curated the IPAPACK, a massively multilingual speech corpora with phonemic transcriptions, encompassing more than 115 languages from diverse language families, selectively checked by linguists. Based on the IPAPACK, we propose CLAP-IPA, a multi-lingual phoneme-speech contrastive embedding model capable of open-vocabulary matching between arbitrary speech signals and phonemic sequences. The proposed model was tested on 95 unseen languages, showing strong generalizability across languages. Temporal alignments between phonemes and speech signals also emerged from contrastive training, enabling zeroshot forced alignment in unseen languages. We further introduced a neural forced aligner IPA-ALIGNER by finetuning CLAP-IPA with the Forward-Sum loss to learn better phone-to-audio alignment. Evaluation results suggest that IPA-ALIGNER can generalize to unseen languages without adaptation.
format Preprint
id arxiv_https___arxiv_org_abs_2311_08323
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language
Zhu, Jian
Yang, Changbing
Samir, Farhan
Islam, Jahurul
Computation and Language
Sound
Audio and Speech Processing
In this project, we demonstrate that phoneme-based models for speech processing can achieve strong crosslinguistic generalizability to unseen languages. We curated the IPAPACK, a massively multilingual speech corpora with phonemic transcriptions, encompassing more than 115 languages from diverse language families, selectively checked by linguists. Based on the IPAPACK, we propose CLAP-IPA, a multi-lingual phoneme-speech contrastive embedding model capable of open-vocabulary matching between arbitrary speech signals and phonemic sequences. The proposed model was tested on 95 unseen languages, showing strong generalizability across languages. Temporal alignments between phonemes and speech signals also emerged from contrastive training, enabling zeroshot forced alignment in unseen languages. We further introduced a neural forced aligner IPA-ALIGNER by finetuning CLAP-IPA with the Forward-Sum loss to learn better phone-to-audio alignment. Evaluation results suggest that IPA-ALIGNER can generalize to unseen languages without adaptation.
title The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2311.08323