Summarizing Speech: A Comprehensive Survey

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Retkowski, Fabian, Züfle, Maike, Sudmann, Andreas, Pfau, Dinah, Watanabe, Shinji, Niehues, Jan, Waibel, Alexander
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915559599243264
author Retkowski, Fabian
Züfle, Maike
Sudmann, Andreas
Pfau, Dinah
Watanabe, Shinji
Niehues, Jan
Waibel, Alexander
author_facet Retkowski, Fabian
Züfle, Maike
Sudmann, Andreas
Pfau, Dinah
Watanabe, Shinji
Niehues, Jan
Waibel, Alexander
contents Speech summarization has become an essential tool for efficiently managing and accessing the growing volume of spoken and audiovisual content. However, despite its increasing importance, speech summarization remains loosely defined. The field intersects with several research areas, including speech recognition, text summarization, and specific applications like meeting summarization. This survey not only examines existing datasets and evaluation protocols, which are crucial for assessing the quality of summarization approaches, but also synthesizes recent developments in the field, highlighting the shift from traditional systems to advanced models like fine-tuned cascaded architectures and end-to-end solutions. In doing so, we surface the ongoing challenges, such as the need for realistic evaluation benchmarks, multilingual datasets, and long-context handling.
format Preprint
id arxiv_https___arxiv_org_abs_2504_08024
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Summarizing Speech: A Comprehensive Survey
Retkowski, Fabian
Züfle, Maike
Sudmann, Andreas
Pfau, Dinah
Watanabe, Shinji
Niehues, Jan
Waibel, Alexander
Computation and Language
Sound
Audio and Speech Processing
Speech summarization has become an essential tool for efficiently managing and accessing the growing volume of spoken and audiovisual content. However, despite its increasing importance, speech summarization remains loosely defined. The field intersects with several research areas, including speech recognition, text summarization, and specific applications like meeting summarization. This survey not only examines existing datasets and evaluation protocols, which are crucial for assessing the quality of summarization approaches, but also synthesizes recent developments in the field, highlighting the shift from traditional systems to advanced models like fine-tuned cascaded architectures and end-to-end solutions. In doing so, we surface the ongoing challenges, such as the need for realistic evaluation benchmarks, multilingual datasets, and long-context handling.
title Summarizing Speech: A Comprehensive Survey
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2504.08024