Beyond the Utterance: An Empirical Study of Very Long Context Speech Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Flynn, Robert, Ragni, Anton |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Much Context Does My Attention-Based ASR System Need?
von: Flynn, Robert, et al.
Veröffentlicht: (2023)
von: Flynn, Robert, et al.
Veröffentlicht: (2023)
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
von: Leung, Wing-Zin, et al.
Veröffentlicht: (2024)
von: Leung, Wing-Zin, et al.
Veröffentlicht: (2024)
Self-Train Before You Transcribe
von: Flynn, Robert, et al.
Veröffentlicht: (2024)
von: Flynn, Robert, et al.
Veröffentlicht: (2024)
Decoding Order Matters in Autoregressive Speech Synthesis
von: Zhao, Minghui, et al.
Veröffentlicht: (2026)
von: Zhao, Minghui, et al.
Veröffentlicht: (2026)
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
von: Yen, Hao, et al.
Veröffentlicht: (2024)
von: Yen, Hao, et al.
Veröffentlicht: (2024)
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
Multi-Utterance Speech Separation and Association Trained on Short Segments
von: Wang, Yuzhu, et al.
Veröffentlicht: (2025)
von: Wang, Yuzhu, et al.
Veröffentlicht: (2025)
Attractor-Based Speech Separation of Multiple Utterances by Unknown Number of Speakers
von: Wang, Yuzhu, et al.
Veröffentlicht: (2025)
von: Wang, Yuzhu, et al.
Veröffentlicht: (2025)
Improving Speech-based Emotion Recognition with Contextual Utterance Analysis and LLMs
von: Zhang, Enshi, et al.
Veröffentlicht: (2024)
von: Zhang, Enshi, et al.
Veröffentlicht: (2024)
Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems
von: Zink, Oswald, et al.
Veröffentlicht: (2024)
von: Zink, Oswald, et al.
Veröffentlicht: (2024)
Long-Context Speech Synthesis with Context-Aware Memory
von: Li, Zhipeng, et al.
Veröffentlicht: (2025)
von: Li, Zhipeng, et al.
Veröffentlicht: (2025)
Score-Based Training for Energy-Based TTS Models
von: Sun, Wanli, et al.
Veröffentlicht: (2025)
von: Sun, Wanli, et al.
Veröffentlicht: (2025)
LCB-net: Long-Context Biasing for Audio-Visual Speech Recognition
von: Yu, Fan, et al.
Veröffentlicht: (2024)
von: Yu, Fan, et al.
Veröffentlicht: (2024)
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
von: Bai, Ye, et al.
Veröffentlicht: (2024)
von: Bai, Ye, et al.
Veröffentlicht: (2024)
Cross-Utterance Conditioned VAE for Speech Generation
von: Li, Yang, et al.
Veröffentlicht: (2023)
von: Li, Yang, et al.
Veröffentlicht: (2023)
Seeing the Context: Rich Visual Context-Aware Speech Recognition via Multimodal Reasoning
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
Advanced Long-Content Speech Recognition With Factorized Neural Transducer
von: Gong, Xun, et al.
Veröffentlicht: (2024)
von: Gong, Xun, et al.
Veröffentlicht: (2024)
Improving Short Utterance Anti-Spoofing with AASIST2
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2023)
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2023)
Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired Users using Intermediate ASR Features and Human Memory Models
von: Mogridge, Rhiannon, et al.
Veröffentlicht: (2024)
von: Mogridge, Rhiannon, et al.
Veröffentlicht: (2024)
A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition
von: Wang, Shiyao, et al.
Veröffentlicht: (2025)
von: Wang, Shiyao, et al.
Veröffentlicht: (2025)
Efficient and Robust Long-Form Speech Recognition with Hybrid H3-Conformer
von: Honda, Tomoki, et al.
Veröffentlicht: (2024)
von: Honda, Tomoki, et al.
Veröffentlicht: (2024)
The Faetar Benchmark: Speech Recognition in a Very Under-Resourced Language
von: Ong, Michael, et al.
Veröffentlicht: (2024)
von: Ong, Michael, et al.
Veröffentlicht: (2024)
Improving Streaming Speech Recognition With Time-Shifted Contextual Attention And Dynamic Right Context Masking
von: Le, Khanh, et al.
Veröffentlicht: (2025)
von: Le, Khanh, et al.
Veröffentlicht: (2025)
In-Materia Speech Recognition
von: Zolfagharinejad, Mohamadreza, et al.
Veröffentlicht: (2024)
von: Zolfagharinejad, Mohamadreza, et al.
Veröffentlicht: (2024)
Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2026)
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2026)
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
von: Chen, Peikun, et al.
Veröffentlicht: (2024)
von: Chen, Peikun, et al.
Veröffentlicht: (2024)
An Empirical Study on the Impact of Positional Encoding in Transformer-based Monaural Speech Enhancement
von: Zhang, Qiquan, et al.
Veröffentlicht: (2024)
von: Zhang, Qiquan, et al.
Veröffentlicht: (2024)
Beyond Performance Plateaus: A Comprehensive Study on Scalability in Speech Enhancement
von: Zhang, Wangyou, et al.
Veröffentlicht: (2024)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2024)
End-to-End Target Speaker Speech Recognition Using Context-Aware Attention Mechanisms for Challenging Enrollment Scenario
von: Ghane, Mohsen, et al.
Veröffentlicht: (2025)
von: Ghane, Mohsen, et al.
Veröffentlicht: (2025)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2023)
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2023)
Robust Speech Recognition with Schrödinger Bridge-Based Speech Enhancement
von: Nasretdinov, Rauf, et al.
Veröffentlicht: (2025)
von: Nasretdinov, Rauf, et al.
Veröffentlicht: (2025)
Speech Emotion Recognition with ASR Integration
von: Li, Yuanchao
Veröffentlicht: (2026)
von: Li, Yuanchao
Veröffentlicht: (2026)
Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
von: Alsayegh, Ali, et al.
Veröffentlicht: (2025)
von: Alsayegh, Ali, et al.
Veröffentlicht: (2025)
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
von: Tian, Jingguang, et al.
Veröffentlicht: (2024)
von: Tian, Jingguang, et al.
Veröffentlicht: (2024)
Exploring the Integration of Large Language Models into Automatic Speech Recognition Systems: An Empirical Study
von: Min, Zeping, et al.
Veröffentlicht: (2023)
von: Min, Zeping, et al.
Veröffentlicht: (2023)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
Mamba-based Decoder-Only Approach with Bidirectional Speech Modeling for Speech Recognition
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2024)
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2024)
Speech Recognition for Analysis of Police Radio Communication
von: Srivastava, Tejes, et al.
Veröffentlicht: (2024)
von: Srivastava, Tejes, et al.
Veröffentlicht: (2024)
Two-pass Endpoint Detection for Speech Recognition
von: Raju, Anirudh, et al.
Veröffentlicht: (2024)
von: Raju, Anirudh, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
How Much Context Does My Attention-Based ASR System Need?
von: Flynn, Robert, et al.
Veröffentlicht: (2023) -
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
von: Leung, Wing-Zin, et al.
Veröffentlicht: (2024) -
Self-Train Before You Transcribe
von: Flynn, Robert, et al.
Veröffentlicht: (2024) -
Decoding Order Matters in Autoregressive Speech Synthesis
von: Zhao, Minghui, et al.
Veröffentlicht: (2026) -
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
von: Yen, Hao, et al.
Veröffentlicht: (2024)