Adaptive Endpointing with Deep Contextual Multi-armed Bandits
Fuente:
arXiv
Saved in:
| Main Authors: | Min, Do June, Stolcke, Andreas, Raju, Anirudh, Vaz, Colin, He, Di, Ravichandran, Venkatesh, Trinh, Viet Anh |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Two-pass Endpoint Detection for Speech Recognition
by: Raju, Anirudh, et al.
Published: (2024)
by: Raju, Anirudh, et al.
Published: (2024)
Reducing Geographic Disparities in Automatic Speech Recognition via Elastic Weight Consolidation
by: Trinh, Viet Anh, et al.
Published: (2022)
by: Trinh, Viet Anh, et al.
Published: (2022)
Turn-taking and Backchannel Prediction with Acoustic and Large Language Model Fusion
by: Wang, Jinhan, et al.
Published: (2024)
by: Wang, Jinhan, et al.
Published: (2024)
Cross-utterance ASR Rescoring with Graph-based Label Propagation
by: Tankasala, Srinath, et al.
Published: (2023)
by: Tankasala, Srinath, et al.
Published: (2023)
Toward Fairness in Speech Recognition: Discovery and mitigation of performance disparities
by: Dheram, Pranav, et al.
Published: (2022)
by: Dheram, Pranav, et al.
Published: (2022)
MLLM-based Speech Recognition: When and How is Multimodality Beneficial?
by: Guan, Yiwen, et al.
Published: (2025)
by: Guan, Yiwen, et al.
Published: (2025)
Multimodal Attention Merging for Improved Speech Recognition and Audio Event Classification
by: Sundar, Anirudh S., et al.
Published: (2023)
by: Sundar, Anirudh S., et al.
Published: (2023)
Discrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing
by: Trinh, Viet Anh, et al.
Published: (2024)
by: Trinh, Viet Anh, et al.
Published: (2024)
Learning When to Trust Which Teacher for Weakly Supervised ASR
by: Agrawal, Aakriti, et al.
Published: (2023)
by: Agrawal, Aakriti, et al.
Published: (2023)
STraDa: A Singer Traits Dataset
by: Kong, Yuexuan, et al.
Published: (2024)
by: Kong, Yuexuan, et al.
Published: (2024)
openFEAT: Improving Speaker Identification by Open-set Few-shot Embedding Adaptation with Transformer
by: C, Kishan K, et al.
Published: (2022)
by: C, Kishan K, et al.
Published: (2022)
Multi-modal Speech Transformer Decoders: When Do Multiple Modalities Improve Accuracy?
by: Guan, Yiwen, et al.
Published: (2024)
by: Guan, Yiwen, et al.
Published: (2024)
Streaming Speech-to-Confusion Network Speech Recognition
by: Filimonov, Denis, et al.
Published: (2023)
by: Filimonov, Denis, et al.
Published: (2023)
Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
by: He, Xinlu, et al.
Published: (2025)
by: He, Xinlu, et al.
Published: (2025)
Improving fairness in speaker verification via Group-adapted Fusion Network
by: Shen, Hua, et al.
Published: (2022)
by: Shen, Hua, et al.
Published: (2022)
Adversarial Reweighting for Speaker Verification Fairness
by: Jin, Minho, et al.
Published: (2022)
by: Jin, Minho, et al.
Published: (2022)
Enhancing Conversational TTS with Cascaded Prompting and ICL-Based Online Reinforcement Learning
by: Ouyang, Zhicheng, et al.
Published: (2026)
by: Ouyang, Zhicheng, et al.
Published: (2026)
Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction
by: Wu, Weijie, et al.
Published: (2025)
by: Wu, Weijie, et al.
Published: (2025)
Post-Training Embedding Alignment for Decoupling Enrollment and Runtime Speaker Recognition Models
by: Gao, Chenyang, et al.
Published: (2024)
by: Gao, Chenyang, et al.
Published: (2024)
AdaPTwin: Low-Cost Adaptive Compression of Product Twins in Transformers
by: Biju, Emil, et al.
Published: (2024)
by: Biju, Emil, et al.
Published: (2024)
PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers
by: Pandey, Rahul, et al.
Published: (2023)
by: Pandey, Rahul, et al.
Published: (2023)
Deep CLAS: Deep Contextual Listen, Attend and Spell
by: Wang, Mengzhi, et al.
Published: (2024)
by: Wang, Mengzhi, et al.
Published: (2024)
Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM
by: Prakash, Jeena, et al.
Published: (2025)
by: Prakash, Jeena, et al.
Published: (2025)
Deep learning classification system for coconut maturity levels based on acoustic signals
by: Caladcad, June Anne, et al.
Published: (2024)
by: Caladcad, June Anne, et al.
Published: (2024)
Speech Retrieval-Augmented Generation without Automatic Speech Recognition
by: Min, Do June, et al.
Published: (2024)
by: Min, Do June, et al.
Published: (2024)
Sound Field Estimation Using Deep Kernel Learning Regularized by the Wave Equation
by: Sundström, David, et al.
Published: (2024)
by: Sundström, David, et al.
Published: (2024)
Linearly Constrained Deep Beamformer for Multi-Speaker Scenarios
by: Zaidel, Ilai, et al.
Published: (2026)
by: Zaidel, Ilai, et al.
Published: (2026)
Contextual Biasing for Streaming ASR via CTC-based Word Spotting
by: Tsai, Kai-Chen, et al.
Published: (2026)
by: Tsai, Kai-Chen, et al.
Published: (2026)
Effective Modeling of Critical Contextual Information for TDNN-based Speaker Verification
by: Weng, Shilong, et al.
Published: (2025)
by: Weng, Shilong, et al.
Published: (2025)
Contextual Biasing for LLM-Based ASR with Hotword Retrieval and Reinforcement Learning
by: Kong, YuXiang, et al.
Published: (2025)
by: Kong, YuXiang, et al.
Published: (2025)
Towards High-Fidelity and Controllable Bioacoustic Generation via Enhanced Diffusion Learning
by: Song, Tianyu, et al.
Published: (2025)
by: Song, Tianyu, et al.
Published: (2025)
Improving speaker verification robustness with synthetic emotional utterances
by: Koditala, Nikhil Kumar, et al.
Published: (2024)
by: Koditala, Nikhil Kumar, et al.
Published: (2024)
Comprehensive Audio Query Handling System with Integrated Expert Models and Contextual Understanding
by: Naveen, Vakada, et al.
Published: (2024)
by: Naveen, Vakada, et al.
Published: (2024)
Target Speaker Extraction by Directly Exploiting Contextual Information in the Time-Frequency Domain
by: Yang, Xue, et al.
Published: (2024)
by: Yang, Xue, et al.
Published: (2024)
RLBR: Reinforcement Learning with Biasing Rewards for Contextual Speech Large Language Models
by: Ren, Bo, et al.
Published: (2026)
by: Ren, Bo, et al.
Published: (2026)
JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions
by: Zhang, Leying, et al.
Published: (2026)
by: Zhang, Leying, et al.
Published: (2026)
A Multi-Channel Auditory Signal Encoder with Adaptive Resolution Using Volatile Memristors
by: Guo, Dongxu, et al.
Published: (2025)
by: Guo, Dongxu, et al.
Published: (2025)
Unifying Streaming and Non-streaming Zipformer-based ASR
by: Sharma, Bidisha, et al.
Published: (2025)
by: Sharma, Bidisha, et al.
Published: (2025)
Empowering Multimodal Respiratory Sound Classification with Counterfactual Adversarial Debiasing for Out-of-Distribution Robustness
by: Koo, Heejoon, et al.
Published: (2025)
by: Koo, Heejoon, et al.
Published: (2025)
Contextual Biasing for ASR in Speech LLM with Common Word Cues and Bias Word Position Prediction
by: Novitasari, Sashi, et al.
Published: (2026)
by: Novitasari, Sashi, et al.
Published: (2026)
Similar Items
-
Two-pass Endpoint Detection for Speech Recognition
by: Raju, Anirudh, et al.
Published: (2024) -
Reducing Geographic Disparities in Automatic Speech Recognition via Elastic Weight Consolidation
by: Trinh, Viet Anh, et al.
Published: (2022) -
Turn-taking and Backchannel Prediction with Acoustic and Large Language Model Fusion
by: Wang, Jinhan, et al.
Published: (2024) -
Cross-utterance ASR Rescoring with Graph-based Label Propagation
by: Tankasala, Srinath, et al.
Published: (2023) -
Toward Fairness in Speech Recognition: Discovery and mitigation of performance disparities
by: Dheram, Pranav, et al.
Published: (2022)