Improving ASR Contextual Biasing with Guided Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tang, Jiyang, Kim, Kwangyoun, Shon, Suwon, Wu, Felix, Sridhar, Prashant, Watanabe, Shinji |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding
von: Shon, Suwon, et al.
Veröffentlicht: (2024)
von: Shon, Suwon, et al.
Veröffentlicht: (2024)
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
Lightweight Prompt Biasing for Contextualized End-to-End ASR Systems
von: Ren, Bo, et al.
Veröffentlicht: (2025)
von: Ren, Bo, et al.
Veröffentlicht: (2025)
CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2026)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2026)
Improving Multilingual ASR in the Wild Using Simple N-best Re-ranking
von: Yan, Brian, et al.
Veröffentlicht: (2024)
von: Yan, Brian, et al.
Veröffentlicht: (2024)
SSVD: Structured SVD for Parameter-Efficient Fine-Tuning and Benchmarking under Domain Shift in ASR
von: Wang, Pu, et al.
Veröffentlicht: (2025)
von: Wang, Pu, et al.
Veröffentlicht: (2025)
WCTC-Biasing: Retraining-free Contextual Biasing ASR with Wildcard CTC-based Keyword Spotting and Inter-layer Biasing
von: Nakagome, Yu, et al.
Veröffentlicht: (2025)
von: Nakagome, Yu, et al.
Veröffentlicht: (2025)
Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator
von: Sun, Guangzhi, et al.
Veröffentlicht: (2022)
von: Sun, Guangzhi, et al.
Veröffentlicht: (2022)
Semi-Autoregressive Streaming ASR With Label Context
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2025)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2025)
Who Spoke What When? Evaluating Spoken Language Models for Conversational ASR with Semantic and Overlap-Aware Metrics
von: Tawara, Naohiro, et al.
Veröffentlicht: (2026)
von: Tawara, Naohiro, et al.
Veröffentlicht: (2026)
Enhancing the Robustness of Contextual ASR to Varying Biasing Information Volumes Through Purified Semantic Correlation Joint Modeling
von: Gu, Yue, et al.
Veröffentlicht: (2025)
von: Gu, Yue, et al.
Veröffentlicht: (2025)
Improving Neural Biasing for Contextual Speech Recognition by Early Context Injection and Text Perturbation
von: Huang, Ruizhe, et al.
Veröffentlicht: (2024)
von: Huang, Ruizhe, et al.
Veröffentlicht: (2024)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
Automatic Text Pronunciation Correlation Generation and Application for Contextual Biasing
von: Cheng, Gaofeng, et al.
Veröffentlicht: (2025)
von: Cheng, Gaofeng, et al.
Veröffentlicht: (2025)
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
von: Sharma, Roshan, et al.
Veröffentlicht: (2024)
von: Sharma, Roshan, et al.
Veröffentlicht: (2024)
DYNAC: Dynamic Vocabulary based Non-Autoregressive Contextualization for Speech Recognition
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
Contextual Biasing for LLM-Based ASR with Hotword Retrieval and Reinforcement Learning
von: Kong, YuXiang, et al.
Veröffentlicht: (2025)
von: Kong, YuXiang, et al.
Veröffentlicht: (2025)
The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties
von: Chen, William, et al.
Veröffentlicht: (2025)
von: Chen, William, et al.
Veröffentlicht: (2025)
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
von: Wang, He, et al.
Veröffentlicht: (2025)
von: Wang, He, et al.
Veröffentlicht: (2025)
CTC-Assisted LLM-Based Contextual ASR
von: Yang, Guanrou, et al.
Veröffentlicht: (2024)
von: Yang, Guanrou, et al.
Veröffentlicht: (2024)
Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC
von: Wang, Qingzheng, et al.
Veröffentlicht: (2025)
von: Wang, Qingzheng, et al.
Veröffentlicht: (2025)
ASR Under the Stethoscope: Evaluating Biases in Clinical Speech Recognition across Indian Languages
von: Kumar, Subham, et al.
Veröffentlicht: (2025)
von: Kumar, Subham, et al.
Veröffentlicht: (2025)
Joint Beam Search Integrating CTC, Attention, and Transducer Decoders
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
Contextual Biasing for Streaming ASR via CTC-based Word Spotting
von: Tsai, Kai-Chen, et al.
Veröffentlicht: (2026)
von: Tsai, Kai-Chen, et al.
Veröffentlicht: (2026)
Finding Task-specific Subnetworks in Multi-task Spoken Language Understanding Model
von: Futami, Hayato, et al.
Veröffentlicht: (2024)
von: Futami, Hayato, et al.
Veröffentlicht: (2024)
Decoder-only Architecture for Streaming End-to-end Speech Recognition
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024)
AutoMode-ASR: Learning to Select ASR Systems for Better Quality and Cost
von: Gündüz, Ahmet, et al.
Veröffentlicht: (2024)
von: Gündüz, Ahmet, et al.
Veröffentlicht: (2024)
Sometin Beta Pass Notin (SBPN): Improving Multilingual ASR for Nigerian Languages via Knowledge Distillation
von: Ogun, Sewade
Veröffentlicht: (2026)
von: Ogun, Sewade
Veröffentlicht: (2026)
Large Language Models based ASR Error Correction for Child Conversations
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
A Practitioner's Guide to Building ASR Models for Low-Resource Languages: A Case Study on Scottish Gaelic
von: Klejch, Ondřej, et al.
Veröffentlicht: (2025)
von: Klejch, Ondřej, et al.
Veröffentlicht: (2025)
How Much Context Does My Attention-Based ASR System Need?
von: Flynn, Robert, et al.
Veröffentlicht: (2023)
von: Flynn, Robert, et al.
Veröffentlicht: (2023)
PromptASR for contextualized ASR with controllable style
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
Streaming Bilingual End-to-End ASR model using Attention over Multiple Softmax
von: Patil, Aditya, et al.
Veröffentlicht: (2024)
von: Patil, Aditya, et al.
Veröffentlicht: (2024)
Selective Attention Merging for low resource tasks: A case study of Child ASR
von: Shankar, Natarajan Balaji, et al.
Veröffentlicht: (2025)
von: Shankar, Natarajan Balaji, et al.
Veröffentlicht: (2025)
InterBiasing: Boost Unseen Word Recognition through Biasing Intermediate Predictions
von: Nakagome, Yu, et al.
Veröffentlicht: (2024)
von: Nakagome, Yu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding
von: Shon, Suwon, et al.
Veröffentlicht: (2024) -
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024) -
Lightweight Prompt Biasing for Contextualized End-to-End ASR Systems
von: Ren, Bo, et al.
Veröffentlicht: (2025) -
CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2026) -
Improving Multilingual ASR in the Wild Using Simple N-best Re-ranking
von: Yan, Brian, et al.
Veröffentlicht: (2024)