From Text Segmentation to Smart Chaptering: A Novel Benchmark for Structuring Video Transcriptions
Fuente:
arXiv
Saved in:
| Main Authors: | Retkowski, Fabian, Waibel, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Paragraph Segmentation Revisited: Towards a Standard Task for Structuring Speech
by: Retkowski, Fabian, et al.
Published: (2025)
by: Retkowski, Fabian, et al.
Published: (2025)
Zero-Shot Strategies for Length-Controllable Summarization
by: Retkowski, Fabian, et al.
Published: (2024)
by: Retkowski, Fabian, et al.
Published: (2024)
Beyond Transcripts: A Renewed Perspective on Audio Chaptering
by: Retkowski, Fabian, et al.
Published: (2026)
by: Retkowski, Fabian, et al.
Published: (2026)
The AI Co-Ethnographer: How Far Can Automation Take Qualitative Research?
by: Retkowski, Fabian, et al.
Published: (2025)
by: Retkowski, Fabian, et al.
Published: (2025)
Titanic Calling: Low Bandwidth Video Conference from the Titanic Wreck
by: Eyiokur, Fevziye Irem, et al.
Published: (2024)
by: Eyiokur, Fevziye Irem, et al.
Published: (2024)
Do What I Say: A Spoken Prompt Dataset for Instruction-Following
by: Züfle, Maike, et al.
Published: (2026)
by: Züfle, Maike, et al.
Published: (2026)
Summarizing Speech: A Comprehensive Survey
by: Retkowski, Fabian, et al.
Published: (2025)
by: Retkowski, Fabian, et al.
Published: (2025)
BOOM: Beyond Only One Modality KIT's Multimodal Multilingual Lecture Companion
by: Koneru, Sai, et al.
Published: (2025)
by: Koneru, Sai, et al.
Published: (2025)
Lombard Speech Synthesis for Any Voice with Controllable Style Embeddings
by: Akti, Seymanur, et al.
Published: (2026)
by: Akti, Seymanur, et al.
Published: (2026)
Continuously Learning New Words in Automatic Speech Recognition
by: Huber, Christian, et al.
Published: (2024)
by: Huber, Christian, et al.
Published: (2024)
Context Biasing for Pronunciation-Orthography Mismatch in Automatic Speech Recognition
by: Huber, Christian, et al.
Published: (2025)
by: Huber, Christian, et al.
Published: (2025)
MUSCAT: MUltilingual, SCientific ConversATion Benchmark
by: Sinhamahapatra, Supriti, et al.
Published: (2026)
by: Sinhamahapatra, Supriti, et al.
Published: (2026)
A Cocktail-Party Benchmark: Multi-Modal dataset and Comparative Evaluation Results
by: Nguyen, Thai-Binh, et al.
Published: (2025)
by: Nguyen, Thai-Binh, et al.
Published: (2025)
Convoifilter: A case study of doing cocktail party speech recognition
by: Nguyen, Thai-Binh, et al.
Published: (2023)
by: Nguyen, Thai-Binh, et al.
Published: (2023)
Handling Numeric Expressions in Automatic Speech Recognition
by: Huber, Christian, et al.
Published: (2024)
by: Huber, Christian, et al.
Published: (2024)
Decoupled Vocabulary Learning Enables Zero-Shot Translation from Unseen Languages
by: Mullov, Carlos, et al.
Published: (2024)
by: Mullov, Carlos, et al.
Published: (2024)
IPA Transcription of Bengali Texts
by: Fatema, Kanij, et al.
Published: (2024)
by: Fatema, Kanij, et al.
Published: (2024)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
by: Nguyen, Thai-Binh, et al.
Published: (2024)
by: Nguyen, Thai-Binh, et al.
Published: (2024)
Lyrics Transcription for Humans: A Readability-Aware Benchmark
by: Cífka, Ondřej, et al.
Published: (2024)
by: Cífka, Ondřej, et al.
Published: (2024)
End-to-End Evaluation for Low-Latency Simultaneous Speech Translation
by: Huber, Christian, et al.
Published: (2023)
by: Huber, Christian, et al.
Published: (2023)
PIER: A Novel Metric for Evaluating What Matters in Code-Switching
by: Ugan, Enes Yavuz, et al.
Published: (2025)
by: Ugan, Enes Yavuz, et al.
Published: (2025)
Distinguishing Repetition Disfluency from Morphological Reduplication in Bangla ASR Transcripts: A Novel Corpus and Benchmarking Analysis
by: Arpa, Zaara Zabeen, et al.
Published: (2025)
by: Arpa, Zaara Zabeen, et al.
Published: (2025)
Cocktail-Party Audio-Visual Speech Recognition
by: Nguyen, Thai-Binh, et al.
Published: (2025)
by: Nguyen, Thai-Binh, et al.
Published: (2025)
Investigating Transcription Normalization in the Faetar ASR Benchmark
by: Peckham, Leo, et al.
Published: (2025)
by: Peckham, Leo, et al.
Published: (2025)
TreeSeg: Hierarchical Topic Segmentation of Large Transcripts
by: Gklezakos, Dimitrios C., et al.
Published: (2024)
by: Gklezakos, Dimitrios C., et al.
Published: (2024)
Evaluating Text Style Transfer: A Nine-Language Benchmark for Text Detoxification
by: Protasov, Vitaly, et al.
Published: (2025)
by: Protasov, Vitaly, et al.
Published: (2025)
Towards continually learning new languages
by: Pham, Ngoc-Quan, et al.
Published: (2022)
by: Pham, Ngoc-Quan, et al.
Published: (2022)
Streaming Translation and Transcription Through Speech-to-Text Causal Alignment
by: Koshkin, Roman, et al.
Published: (2026)
by: Koshkin, Roman, et al.
Published: (2026)
Accent conversion using discrete units with parallel data synthesized from controllable accented TTS
by: Nguyen, Tuan Nam, et al.
Published: (2024)
by: Nguyen, Tuan Nam, et al.
Published: (2024)
Weight Factorization and Centralization for Continual Learning in Speech Recognition
by: Ugan, Enes Yavuz, et al.
Published: (2025)
by: Ugan, Enes Yavuz, et al.
Published: (2025)
Algorithms For Automatic Accentuation And Transcription Of Russian Texts In Speech Recognition Systems
by: Iakovenko, Olga, et al.
Published: (2024)
by: Iakovenko, Olga, et al.
Published: (2024)
Reversed in Time: A Novel Temporal-Emphasized Benchmark for Cross-Modal Video-Text Retrieval
by: Du, Yang, et al.
Published: (2024)
by: Du, Yang, et al.
Published: (2024)
Mark My Words: A Robust Multilingual Model for Punctuation in Text and Speech Transcripts
by: Pulipaka, Sidharth, et al.
Published: (2025)
by: Pulipaka, Sidharth, et al.
Published: (2025)
BanglaIPA: Towards Robust Text-to-IPA Transcription with Contextual Rewriting in Bengali
by: Hasan, Jakir, et al.
Published: (2026)
by: Hasan, Jakir, et al.
Published: (2026)
VideoAVE: A Multi-Attribute Video-to-Text Attribute Value Extraction Dataset and Benchmark Models
by: Cheng, Ming, et al.
Published: (2025)
by: Cheng, Ming, et al.
Published: (2025)
Identifying Speakers in Dialogue Transcripts: A Text-based Approach Using Pretrained Language Models
by: Nguyen, Minh, et al.
Published: (2024)
by: Nguyen, Minh, et al.
Published: (2024)
BoundRL: Efficient Structured Text Segmentation through Reinforced Boundary Generation
by: Li, Haoyuan, et al.
Published: (2025)
by: Li, Haoyuan, et al.
Published: (2025)
TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
by: Feng, Weixi, et al.
Published: (2024)
by: Feng, Weixi, et al.
Published: (2024)
Bayesian Low-Rank Factorization for Robust Model Adaptation
by: Ugan, Enes Yavuz, et al.
Published: (2025)
by: Ugan, Enes Yavuz, et al.
Published: (2025)
Adapting Language Balance in Code-Switching Speech
by: Ugan, Enes Yavuz, et al.
Published: (2025)
by: Ugan, Enes Yavuz, et al.
Published: (2025)
Similar Items
-
Paragraph Segmentation Revisited: Towards a Standard Task for Structuring Speech
by: Retkowski, Fabian, et al.
Published: (2025) -
Zero-Shot Strategies for Length-Controllable Summarization
by: Retkowski, Fabian, et al.
Published: (2024) -
Beyond Transcripts: A Renewed Perspective on Audio Chaptering
by: Retkowski, Fabian, et al.
Published: (2026) -
The AI Co-Ethnographer: How Far Can Automation Take Qualitative Research?
by: Retkowski, Fabian, et al.
Published: (2025) -
Titanic Calling: Low Bandwidth Video Conference from the Titanic Wreck
by: Eyiokur, Fevziye Irem, et al.
Published: (2024)