Unsupervised Speech Segmentation: A General Approach Using Speech Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Elmakies, Avishai, Abend, Omri, Adi, Yossi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910775688298496
author Elmakies, Avishai
Abend, Omri
Adi, Yossi
author_facet Elmakies, Avishai
Abend, Omri
Adi, Yossi
contents In this paper, we introduce an unsupervised approach for Speech Segmentation, which builds on previously researched approaches, e.g., Speaker Diarization, while being applicable to an inclusive set of acoustic-semantic distinctions, paving a path towards a general Unsupervised Speech Segmentation approach. Unlike traditional speech and audio segmentation, which mainly focuses on spectral changes in the input signal, e.g., phone segmentation, our approach tries to segment the spoken utterance into chunks with differing acoustic-semantic styles, focusing on acoustic-semantic information that does not translate well into text, e.g., emotion or speaker. While most Speech Segmentation tasks only handle one style change, e.g., emotion diarization, our approach tries to handle multiple acoustic-semantic style changes. Leveraging recent advances in Speech Language Models (SLMs), we propose a simple unsupervised method to segment a given speech utterance. We empirically demonstrate the effectiveness of the proposed approach by considering several setups. Results suggest that the proposed method is superior to the evaluated baselines on boundary detection, segment purity, and over-segmentation. Code is available at https://github.com/avishaiElmakies/unsupervised_speech_segmentation_using_slm.
format Preprint
id arxiv_https___arxiv_org_abs_2501_03711
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unsupervised Speech Segmentation: A General Approach Using Speech Language Models
Elmakies, Avishai
Abend, Omri
Adi, Yossi
Computation and Language
Artificial Intelligence
Machine Learning
Sound
Audio and Speech Processing
In this paper, we introduce an unsupervised approach for Speech Segmentation, which builds on previously researched approaches, e.g., Speaker Diarization, while being applicable to an inclusive set of acoustic-semantic distinctions, paving a path towards a general Unsupervised Speech Segmentation approach. Unlike traditional speech and audio segmentation, which mainly focuses on spectral changes in the input signal, e.g., phone segmentation, our approach tries to segment the spoken utterance into chunks with differing acoustic-semantic styles, focusing on acoustic-semantic information that does not translate well into text, e.g., emotion or speaker. While most Speech Segmentation tasks only handle one style change, e.g., emotion diarization, our approach tries to handle multiple acoustic-semantic style changes. Leveraging recent advances in Speech Language Models (SLMs), we propose a simple unsupervised method to segment a given speech utterance. We empirically demonstrate the effectiveness of the proposed approach by considering several setups. Results suggest that the proposed method is superior to the evaluated baselines on boundary detection, segment purity, and over-segmentation. Code is available at https://github.com/avishaiElmakies/unsupervised_speech_segmentation_using_slm.
title Unsupervised Speech Segmentation: A General Approach Using Speech Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2501.03711