Lightweight Audio Segmentation for Long-form Speech Translation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lee, Jaesong, Kim, Soyoon, Kim, Hanbyul, Chung, Joon Son
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914835156959232
author Lee, Jaesong
Kim, Soyoon
Kim, Hanbyul
Chung, Joon Son
author_facet Lee, Jaesong
Kim, Soyoon
Kim, Hanbyul
Chung, Joon Son
contents Speech segmentation is an essential part of speech translation (ST) systems in real-world scenarios. Since most ST models are designed to process speech segments, long-form audio must be partitioned into shorter segments before translation. Recently, data-driven approaches for the speech segmentation task have been developed. Although the approaches improve overall translation quality, a performance gap exists due to a mismatch between the models and ST systems. In addition, the prior works require large self-supervised speech models, which consume significant computational resources. In this work, we propose a segmentation model that achieves better speech translation quality with a small model size. We propose an ASR-with-punctuation task as an effective pre-training strategy for the segmentation model. We also show that proper integration of the speech segmentation model into the underlying ST system is critical to improve overall translation quality at inference time.
format Preprint
id arxiv_https___arxiv_org_abs_2406_10549
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Lightweight Audio Segmentation for Long-form Speech Translation
Lee, Jaesong
Kim, Soyoon
Kim, Hanbyul
Chung, Joon Son
Audio and Speech Processing
Computation and Language
Sound
Speech segmentation is an essential part of speech translation (ST) systems in real-world scenarios. Since most ST models are designed to process speech segments, long-form audio must be partitioned into shorter segments before translation. Recently, data-driven approaches for the speech segmentation task have been developed. Although the approaches improve overall translation quality, a performance gap exists due to a mismatch between the models and ST systems. In addition, the prior works require large self-supervised speech models, which consume significant computational resources. In this work, we propose a segmentation model that achieves better speech translation quality with a small model size. We propose an ASR-with-punctuation task as an effective pre-training strategy for the segmentation model. We also show that proper integration of the speech segmentation model into the underlying ST system is critical to improve overall translation quality at inference time.
title Lightweight Audio Segmentation for Long-form Speech Translation
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2406.10549