How Much Context Does My Attention-Based ASR System Need?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Flynn, Robert, Ragni, Anton |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond the Utterance: An Empirical Study of Very Long Context Speech Recognition
von: Flynn, Robert, et al.
Veröffentlicht: (2026)
von: Flynn, Robert, et al.
Veröffentlicht: (2026)
Self-Train Before You Transcribe
von: Flynn, Robert, et al.
Veröffentlicht: (2024)
von: Flynn, Robert, et al.
Veröffentlicht: (2024)
Semi-Autoregressive Streaming ASR With Label Context
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
ASR Benchmarking: Need for a More Representative Conversational Dataset
von: Maheshwari, Gaurav, et al.
Veröffentlicht: (2024)
von: Maheshwari, Gaurav, et al.
Veröffentlicht: (2024)
Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR
von: Chen, Qian, et al.
Veröffentlicht: (2023)
von: Chen, Qian, et al.
Veröffentlicht: (2023)
AutoMode-ASR: Learning to Select ASR Systems for Better Quality and Cost
von: Gündüz, Ahmet, et al.
Veröffentlicht: (2024)
von: Gündüz, Ahmet, et al.
Veröffentlicht: (2024)
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
von: Wang, He, et al.
Veröffentlicht: (2025)
von: Wang, He, et al.
Veröffentlicht: (2025)
Mind the Gap: Entity-Preserved Context-Aware ASR Structured Transcriptions
von: Altinok, Duygu
Veröffentlicht: (2025)
von: Altinok, Duygu
Veröffentlicht: (2025)
NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR
von: Xie, Yuan, et al.
Veröffentlicht: (2026)
von: Xie, Yuan, et al.
Veröffentlicht: (2026)
Context-Enhanced Granular Edit Representation for Efficient and Accurate ASR Post-editing
von: Vejsiu, Luan, et al.
Veröffentlicht: (2025)
von: Vejsiu, Luan, et al.
Veröffentlicht: (2025)
Quantizing Whisper-small: How design choices affect ASR performance
von: Söhler, Arthur, et al.
Veröffentlicht: (2025)
von: Söhler, Arthur, et al.
Veröffentlicht: (2025)
PromptASR for contextualized ASR with controllable style
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
Streaming Bilingual End-to-End ASR model using Attention over Multiple Softmax
von: Patil, Aditya, et al.
Veröffentlicht: (2024)
von: Patil, Aditya, et al.
Veröffentlicht: (2024)
Selective Attention Merging for low resource tasks: A case study of Child ASR
von: Shankar, Natarajan Balaji, et al.
Veröffentlicht: (2025)
von: Shankar, Natarajan Balaji, et al.
Veröffentlicht: (2025)
Learnable Pulse Accumulation for On-Device Speech Recognition: How Much Attention Do You Need?
von: Shkolnikov, Yakov Pyotr
Veröffentlicht: (2026)
von: Shkolnikov, Yakov Pyotr
Veröffentlicht: (2026)
The THUEE System Description for the IARPA OpenASR21 Challenge
von: Zhao, Jing, et al.
Veröffentlicht: (2022)
von: Zhao, Jing, et al.
Veröffentlicht: (2022)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
Boosting CTC-Based ASR Using LLM-Based Intermediate Loss Regularization
von: Altinok, Duygu
Veröffentlicht: (2025)
von: Altinok, Duygu
Veröffentlicht: (2025)
ASR-FAIRBENCH: Measuring and Benchmarking Equity Across Speech Recognition Systems
von: Rai, Anand, et al.
Veröffentlicht: (2025)
von: Rai, Anand, et al.
Veröffentlicht: (2025)
Evolutionary Prompt Design for LLM-Based Post-ASR Error Correction
von: Sachdev, Rithik, et al.
Veröffentlicht: (2024)
von: Sachdev, Rithik, et al.
Veröffentlicht: (2024)
Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets
von: Geng, Xuelong, et al.
Veröffentlicht: (2024)
von: Geng, Xuelong, et al.
Veröffentlicht: (2024)
ASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error Correction
von: Wei, Victor Junqiu, et al.
Veröffentlicht: (2024)
von: Wei, Victor Junqiu, et al.
Veröffentlicht: (2024)
Romanization Encoding For Multilingual ASR
von: Ding, Wen, et al.
Veröffentlicht: (2024)
von: Ding, Wen, et al.
Veröffentlicht: (2024)
Promptformer: Prompted Conformer Transducer for ASR
von: Duarte-Torres, Sergio, et al.
Veröffentlicht: (2024)
von: Duarte-Torres, Sergio, et al.
Veröffentlicht: (2024)
Qwen3-ASR Technical Report
von: Shi, Xian, et al.
Veröffentlicht: (2026)
von: Shi, Xian, et al.
Veröffentlicht: (2026)
Revisiting Acoustic Features for Robust ASR
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
Word-Level ASR Quality Estimation for Efficient Corpus Sampling and Post-Editing through Analyzing Attentions of a Reference-Free Metric
von: Javadi, Golara, et al.
Veröffentlicht: (2024)
von: Javadi, Golara, et al.
Veröffentlicht: (2024)
Advancing Hearing Assessment: An ASR-Based Frequency-Specific Speech Test for Diagnosing Presbycusis
von: Bleeck, Stefan
Veröffentlicht: (2025)
von: Bleeck, Stefan
Veröffentlicht: (2025)
Score-Based Training for Energy-Based TTS Models
von: Sun, Wanli, et al.
Veröffentlicht: (2025)
von: Sun, Wanli, et al.
Veröffentlicht: (2025)
Custom Data Augmentation for low resource ASR using Bark and Retrieval-Based Voice Conversion
von: Kamble, Anand, et al.
Veröffentlicht: (2023)
von: Kamble, Anand, et al.
Veröffentlicht: (2023)
Exploring SSL Discrete Tokens for Multilingual ASR
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
Configurable Multilingual ASR with Speech Summary Representations
von: Zhu, Harrison, et al.
Veröffentlicht: (2024)
von: Zhu, Harrison, et al.
Veröffentlicht: (2024)
ManWav: The First Manchu ASR Model
von: Seo, Jean, et al.
Veröffentlicht: (2024)
von: Seo, Jean, et al.
Veröffentlicht: (2024)
Mamba for Streaming ASR Combined with Unimodal Aggregation
von: Fang, Ying, et al.
Veröffentlicht: (2024)
von: Fang, Ying, et al.
Veröffentlicht: (2024)
Extending Whisper with prompt tuning to target-speaker ASR
von: Ma, Hao, et al.
Veröffentlicht: (2023)
von: Ma, Hao, et al.
Veröffentlicht: (2023)
Scalable Offline ASR for Command-Style Dictation in Courtrooms
von: Nethil, Kumarmanas, et al.
Veröffentlicht: (2025)
von: Nethil, Kumarmanas, et al.
Veröffentlicht: (2025)
Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2025)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2025)
Performant ASR Models for Medical Entities in Accented Speech
von: Afonja, Tejumade, et al.
Veröffentlicht: (2024)
von: Afonja, Tejumade, et al.
Veröffentlicht: (2024)
Causal Structure Discovery for Error Diagnostics of Children's ASR
von: Singh, Vishwanath Pratap, et al.
Veröffentlicht: (2025)
von: Singh, Vishwanath Pratap, et al.
Veröffentlicht: (2025)
Reverb: Open-Source ASR and Diarization from Rev
von: Bhandari, Nishchal, et al.
Veröffentlicht: (2024)
von: Bhandari, Nishchal, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Beyond the Utterance: An Empirical Study of Very Long Context Speech Recognition
von: Flynn, Robert, et al.
Veröffentlicht: (2026) -
Self-Train Before You Transcribe
von: Flynn, Robert, et al.
Veröffentlicht: (2024) -
Semi-Autoregressive Streaming ASR With Label Context
von: Arora, Siddhant, et al.
Veröffentlicht: (2023) -
ASR Benchmarking: Need for a More Representative Conversational Dataset
von: Maheshwari, Gaurav, et al.
Veröffentlicht: (2024) -
Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR
von: Chen, Qian, et al.
Veröffentlicht: (2023)