Reading Between the Waves: Robust Topic Segmentation Using Inter-Sentence Audio Features
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Freisinger, Steffen, Seeberger, Philipp, Bocklet, Tobias, Riedhammer, Korbinian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Multi-Level Transcript Segmentation: LoRA Fine-Tuning for Table-of-Contents Generation
von: Freisinger, Steffen, et al.
Veröffentlicht: (2026)
von: Freisinger, Steffen, et al.
Veröffentlicht: (2026)
Predicting Upcoming Stuttering Events from Three-Second Audio: Stratified Evaluation Reveals Severity-Selective Precursors, and the Model Deploys Fully On-Device
von: Kozak, Nazar
Veröffentlicht: (2026)
von: Kozak, Nazar
Veröffentlicht: (2026)
ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
von: Kim, Minu, et al.
Veröffentlicht: (2025)
von: Kim, Minu, et al.
Veröffentlicht: (2025)
Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization
von: Alamr, Meshal, et al.
Veröffentlicht: (2026)
von: Alamr, Meshal, et al.
Veröffentlicht: (2026)
Beyond Levenshtein: Leveraging Multiple Algorithms for Robust Word Error Rate Computations And Granular Error Classifications
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2024)
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2024)
Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition
von: Hori, Takaaki, et al.
Veröffentlicht: (2025)
von: Hori, Takaaki, et al.
Veröffentlicht: (2025)
An End-to-End Approach for Korean Wakeword Systems with Speaker Authentication
von: Seo, Geonwoo
Veröffentlicht: (2025)
von: Seo, Geonwoo
Veröffentlicht: (2025)
Splitformer: An improved early-exit architecture for automatic speech recognition on edge devices
von: Lasbordes, Maxence, et al.
Veröffentlicht: (2025)
von: Lasbordes, Maxence, et al.
Veröffentlicht: (2025)
Emotional Voice Messages (EMOVOME) database: emotion recognition in spontaneous voice messages
von: Zaragozá, Lucía Gómez, et al.
Veröffentlicht: (2024)
von: Zaragozá, Lucía Gómez, et al.
Veröffentlicht: (2024)
FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations
von: Lee, Yoonhyung, et al.
Veröffentlicht: (2026)
von: Lee, Yoonhyung, et al.
Veröffentlicht: (2026)
Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTS
von: Dai, Ziqi, et al.
Veröffentlicht: (2025)
von: Dai, Ziqi, et al.
Veröffentlicht: (2025)
An open-source voice type classifier for child-centered daylong recordings
von: Lavechin, Marvin, et al.
Veröffentlicht: (2020)
von: Lavechin, Marvin, et al.
Veröffentlicht: (2020)
Measuring the Accuracy of Automatic Speech Recognition Solutions
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2024)
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2024)
Communication Access Real-Time Translation Through Collaborative Correction of Automatic Speech Recognition
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2025)
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2025)
Fine-Tuning Large Audio-Language Models with LoRA for Precise Temporal Localization of Prolonged Exposure Therapy Elements
von: BN, Suhas, et al.
Veröffentlicht: (2025)
von: BN, Suhas, et al.
Veröffentlicht: (2025)
Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis
von: Kim, Minu, et al.
Veröffentlicht: (2025)
von: Kim, Minu, et al.
Veröffentlicht: (2025)
Beyond Deep Learning: Speech Segmentation and Phone Classification with Neural Assemblies
von: Adelson, Trevor, et al.
Veröffentlicht: (2026)
von: Adelson, Trevor, et al.
Veröffentlicht: (2026)
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)
Quantifying the effect of speech pathology on automatic and human speaker verification
von: Halpern, Bence Mark, et al.
Veröffentlicht: (2024)
von: Halpern, Bence Mark, et al.
Veröffentlicht: (2024)
SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset
von: Gautam, Sushant, et al.
Veröffentlicht: (2024)
von: Gautam, Sushant, et al.
Veröffentlicht: (2024)
Modeling L1 Influence on L2 Pronunciation: An MFCC-Based Framework for Explainable Machine Learning and Pedagogical Feedback
von: Jahanbin, Peyman
Veröffentlicht: (2025)
von: Jahanbin, Peyman
Veröffentlicht: (2025)
Impact of Phonetics on Speaker Identity in Adversarial Voice Attack
von: Dar, Daniyal Kabir, et al.
Veröffentlicht: (2025)
von: Dar, Daniyal Kabir, et al.
Veröffentlicht: (2025)
MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector
von: Costa-jussà, Marta R., et al.
Veröffentlicht: (2024)
von: Costa-jussà, Marta R., et al.
Veröffentlicht: (2024)
Representation Loss Minimization with Randomized Selection Strategy for Efficient Environmental Fake Audio Detection
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
STRUM: A Spectral Transcription and Rhythm Understanding Model for End-to-End Generation of Playable Rhythm-Game Charts
von: Opria, Joshua
Veröffentlicht: (2026)
von: Opria, Joshua
Veröffentlicht: (2026)
Graph Connectionist Temporal Classification for Phoneme Recognition
von: Grafé, Henry, et al.
Veröffentlicht: (2025)
von: Grafé, Henry, et al.
Veröffentlicht: (2025)
Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Detection of Personal Data in Structured Datasets Using a Large Language Model
von: Ntwali, Albert Agisha, et al.
Veröffentlicht: (2025)
von: Ntwali, Albert Agisha, et al.
Veröffentlicht: (2025)
Evaluating ASR Confidence Scores for Automated Error Detection in User-Assisted Correction Interfaces
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2025)
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2025)
SW-ASR: A Context-Aware Hybrid ASR Pipeline for Robust Single Word Speech Recognition
von: Sharma, Manali, et al.
Veröffentlicht: (2026)
von: Sharma, Manali, et al.
Veröffentlicht: (2026)
SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding
von: Bai, Bingsong, et al.
Veröffentlicht: (2025)
von: Bai, Bingsong, et al.
Veröffentlicht: (2025)
U2++ MoE: Scaling 4.7x parameters with minimal impact on RTF
von: Song, Xingchen, et al.
Veröffentlicht: (2024)
von: Song, Xingchen, et al.
Veröffentlicht: (2024)
Exploring rhythm formant analysis for Indic language classification
von: Gogoi, Parismita, et al.
Veröffentlicht: (2024)
von: Gogoi, Parismita, et al.
Veröffentlicht: (2024)
Self-Improvement for Audio Large Language Model using Unlabeled Speech
von: Wang, Shaowen, et al.
Veröffentlicht: (2025)
von: Wang, Shaowen, et al.
Veröffentlicht: (2025)
Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization
von: Wang, Hsuan-Yu, et al.
Veröffentlicht: (2025)
von: Wang, Hsuan-Yu, et al.
Veröffentlicht: (2025)
Less Stress, More Privacy: Stress Detection on Anonymized Speech of Air Traffic Controllers
von: Viswanathan, Janaki, et al.
Veröffentlicht: (2025)
von: Viswanathan, Janaki, et al.
Veröffentlicht: (2025)
Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Multi-View Multi-Task Modeling with Speech Foundation Models for Speech Forensic Tasks
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children
von: Ahn, Taekyung, et al.
Veröffentlicht: (2024)
von: Ahn, Taekyung, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Towards Multi-Level Transcript Segmentation: LoRA Fine-Tuning for Table-of-Contents Generation
von: Freisinger, Steffen, et al.
Veröffentlicht: (2026) -
Predicting Upcoming Stuttering Events from Three-Second Audio: Stratified Evaluation Reveals Severity-Selective Precursors, and the Model Deploys Fully On-Device
von: Kozak, Nazar
Veröffentlicht: (2026) -
ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
von: Kim, Minu, et al.
Veröffentlicht: (2025) -
Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization
von: Alamr, Meshal, et al.
Veröffentlicht: (2026) -
Beyond Levenshtein: Leveraging Multiple Algorithms for Robust Word Error Rate Computations And Granular Error Classifications
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2024)