Optimizing Contextual Speech Recognition Using Vector Quantization for Efficient Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Flemotomos, Nikolaos, Hsiao, Roger, Swietojanski, Pawel, Hori, Takaaki, Can, Dogan, Zhuang, Xiaodan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Segmental Attention Decoding With Long Form Acoustic Encodings
von: Swietojanski, Pawel, et al.
Veröffentlicht: (2025)
von: Swietojanski, Pawel, et al.
Veröffentlicht: (2025)
Optimizing Byte-level Representation for End-to-end ASR
von: Hsiao, Roger, et al.
Veröffentlicht: (2024)
von: Hsiao, Roger, et al.
Veröffentlicht: (2024)
Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition
von: Hori, Takaaki, et al.
Veröffentlicht: (2025)
von: Hori, Takaaki, et al.
Veröffentlicht: (2025)
Autoregressive Speech Synthesis without Vector Quantization
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
von: Shao, Hang, et al.
Veröffentlicht: (2023)
von: Shao, Hang, et al.
Veröffentlicht: (2023)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
PAC: Pronunciation-Aware Contextualized Large Language Model-based Automatic Speech Recognition
von: Fu, Li, et al.
Veröffentlicht: (2025)
von: Fu, Li, et al.
Veröffentlicht: (2025)
Improving Neural Biasing for Contextual Speech Recognition by Early Context Injection and Text Perturbation
von: Huang, Ruizhe, et al.
Veröffentlicht: (2024)
von: Huang, Ruizhe, et al.
Veröffentlicht: (2024)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary Prediction and Activation
von: Lin, Zhennan, et al.
Veröffentlicht: (2025)
von: Lin, Zhennan, et al.
Veröffentlicht: (2025)
Improving Speech-based Emotion Recognition with Contextual Utterance Analysis and LLMs
von: Zhang, Enshi, et al.
Veröffentlicht: (2024)
von: Zhang, Enshi, et al.
Veröffentlicht: (2024)
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
DYNAC: Dynamic Vocabulary based Non-Autoregressive Contextualization for Speech Recognition
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
von: Wang, He, et al.
Veröffentlicht: (2025)
von: Wang, He, et al.
Veröffentlicht: (2025)
Sequential Editing for Lifelong Training of Speech Recognition Models
von: Kulshreshtha, Devang, et al.
Veröffentlicht: (2024)
von: Kulshreshtha, Devang, et al.
Veröffentlicht: (2024)
RAG-Boost: Retrieval-Augmented Generation Enhanced LLM-based Speech Recognition
von: Wang, Pengcheng, et al.
Veröffentlicht: (2025)
von: Wang, Pengcheng, et al.
Veröffentlicht: (2025)
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
DQ-Data2vec: Decoupling Quantization for Multilingual Speech Recognition
von: Shao, Qijie, et al.
Veröffentlicht: (2025)
von: Shao, Qijie, et al.
Veröffentlicht: (2025)
Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation
von: Wei, Kun, et al.
Veröffentlicht: (2023)
von: Wei, Kun, et al.
Veröffentlicht: (2023)
Retrieval-Augmented Speech Recognition Approach for Domain Challenges
von: Shen, Peng, et al.
Veröffentlicht: (2025)
von: Shen, Peng, et al.
Veröffentlicht: (2025)
Enhancing Large Language Model-based Speech Recognition by Contextualization for Rare and Ambiguous Words
von: Nozawa, Kento, et al.
Veröffentlicht: (2024)
von: Nozawa, Kento, et al.
Veröffentlicht: (2024)
Enhancing Dialogue Speech Recognition with Robust Contextual Awareness via Noise Representation Learning
von: Lee, Wonjun, et al.
Veröffentlicht: (2024)
von: Lee, Wonjun, et al.
Veröffentlicht: (2024)
Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
Streaming Speech-to-Confusion Network Speech Recognition
von: Filimonov, Denis, et al.
Veröffentlicht: (2023)
von: Filimonov, Denis, et al.
Veröffentlicht: (2023)
Speech Retrieval-Augmented Generation without Automatic Speech Recognition
von: Min, Do June, et al.
Veröffentlicht: (2024)
von: Min, Do June, et al.
Veröffentlicht: (2024)
Steering Language Model to Stable Speech Emotion Recognition via Contextual Perception and Chain of Thought
von: Zhao, Zhixian, et al.
Veröffentlicht: (2025)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2025)
Enhancing Fully Formatted End-to-End Speech Recognition with Knowledge Distillation via Multi-Codebook Vector Quantization
von: You, Jian, et al.
Veröffentlicht: (2025)
von: You, Jian, et al.
Veröffentlicht: (2025)
BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2025)
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2025)
Inappropriate Pause Detection In Dysarthric Speech Using Large-Scale Speech Recognition
von: Lee, Jeehyun, et al.
Veröffentlicht: (2024)
von: Lee, Jeehyun, et al.
Veröffentlicht: (2024)
Pheme: Efficient and Conversational Speech Generation
von: Budzianowski, Paweł, et al.
Veröffentlicht: (2024)
von: Budzianowski, Paweł, et al.
Veröffentlicht: (2024)
HENT-SRT: Hierarchical Efficient Neural Transducer with Self-Distillation for Joint Speech Recognition and Translation
von: Hussein, Amir, et al.
Veröffentlicht: (2025)
von: Hussein, Amir, et al.
Veröffentlicht: (2025)
Improved Contextual Recognition In Automatic Speech Recognition Systems By Semantic Lattice Rescoring
von: Sudarshan, Ankitha, et al.
Veröffentlicht: (2023)
von: Sudarshan, Ankitha, et al.
Veröffentlicht: (2023)
ESC: Efficient Speech Coding with Cross-Scale Residual Vector Quantized Transformers
von: Gu, Yuzhe, et al.
Veröffentlicht: (2024)
von: Gu, Yuzhe, et al.
Veröffentlicht: (2024)
SpeechCaps: Advancing Instruction-Based Universal Speech Models with Multi-Talker Speaking Style Captioning
von: Huang, Chien-yu, et al.
Veröffentlicht: (2024)
von: Huang, Chien-yu, et al.
Veröffentlicht: (2024)
SpeechColab Leaderboard: An Open-Source Platform for Automatic Speech Recognition Evaluation
von: Du, Jiayu, et al.
Veröffentlicht: (2024)
von: Du, Jiayu, et al.
Veröffentlicht: (2024)
PhoWhisper: Automatic Speech Recognition for Vietnamese
von: Le, Thanh-Thien, et al.
Veröffentlicht: (2024)
von: Le, Thanh-Thien, et al.
Veröffentlicht: (2024)
Zero-resource Speech Translation and Recognition with LLMs
von: Mundnich, Karel, et al.
Veröffentlicht: (2024)
von: Mundnich, Karel, et al.
Veröffentlicht: (2024)
Speech Recognition Model Improves Text-to-Speech Synthesis using Fine-Grained Reward
von: Wang, Guansu, et al.
Veröffentlicht: (2025)
von: Wang, Guansu, et al.
Veröffentlicht: (2025)
Exploring Generative Error Correction for Dysarthric Speech Recognition
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025)
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025)
Investigating the Impact of Word Informativeness on Speech Emotion Recognition
von: Kakouros, Sofoklis
Veröffentlicht: (2025)
von: Kakouros, Sofoklis
Veröffentlicht: (2025)
SALSA: Speech Aware LLM Adaptation via Learned Steering Activation Vectors
von: Yegorova, Yekaterina, et al.
Veröffentlicht: (2026)
von: Yegorova, Yekaterina, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Segmental Attention Decoding With Long Form Acoustic Encodings
von: Swietojanski, Pawel, et al.
Veröffentlicht: (2025) -
Optimizing Byte-level Representation for End-to-end ASR
von: Hsiao, Roger, et al.
Veröffentlicht: (2024) -
Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition
von: Hori, Takaaki, et al.
Veröffentlicht: (2025) -
Autoregressive Speech Synthesis without Vector Quantization
von: Meng, Lingwei, et al.
Veröffentlicht: (2024) -
DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
von: Shao, Hang, et al.
Veröffentlicht: (2023)