Beyond Transcripts: A Renewed Perspective on Audio Chaptering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Retkowski, Fabian, Züfle, Maike, Nguyen, Thai Binh, Niehues, Jan, Waibel, Alexander
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914613707145216
author Retkowski, Fabian
Züfle, Maike
Nguyen, Thai Binh
Niehues, Jan
Waibel, Alexander
author_facet Retkowski, Fabian
Züfle, Maike
Nguyen, Thai Binh
Niehues, Jan
Waibel, Alexander
contents Audio chaptering, the task of segmenting long-form audio into coherent sections, is increasingly important for navigating podcasts, lectures, and videos. Despite its relevance, research remains limited and text-based, leaving key questions unresolved about leveraging audio information, handling ASR errors, and transcript-free evaluation. We address these gaps through three contributions: (1) a systematic comparison between text-based models with acoustic features, a novel audio-only architecture (AudioSeg) operating on learned audio representations, and multimodal LLMs; (2) empirical analysis of factors affecting performance, including transcript quality, acoustic features, duration, and speaker composition; and (3) formalized evaluation protocols contrasting transcript-dependent text-space protocols with transcript-invariant time-space protocols. Our experiments on YTSeg reveal that AudioSeg substantially outperforms text-based approaches, pauses provide the largest acoustic gains, and MLLMs remain limited by context length and weak instruction following, yet MLLMs are promising on shorter audio.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08979
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond Transcripts: A Renewed Perspective on Audio Chaptering
Retkowski, Fabian
Züfle, Maike
Nguyen, Thai Binh
Niehues, Jan
Waibel, Alexander
Sound
Computation and Language
Audio chaptering, the task of segmenting long-form audio into coherent sections, is increasingly important for navigating podcasts, lectures, and videos. Despite its relevance, research remains limited and text-based, leaving key questions unresolved about leveraging audio information, handling ASR errors, and transcript-free evaluation. We address these gaps through three contributions: (1) a systematic comparison between text-based models with acoustic features, a novel audio-only architecture (AudioSeg) operating on learned audio representations, and multimodal LLMs; (2) empirical analysis of factors affecting performance, including transcript quality, acoustic features, duration, and speaker composition; and (3) formalized evaluation protocols contrasting transcript-dependent text-space protocols with transcript-invariant time-space protocols. Our experiments on YTSeg reveal that AudioSeg substantially outperforms text-based approaches, pauses provide the largest acoustic gains, and MLLMs remain limited by context length and weak instruction following, yet MLLMs are promising on shorter audio.
title Beyond Transcripts: A Renewed Perspective on Audio Chaptering
topic Sound
Computation and Language
url https://arxiv.org/abs/2602.08979