Keyword-Guided Adaptation of Automatic Speech Recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Shamsian, Aviv, Navon, Aviv, Glazer, Neta, Hetz, Gill, Keshet, Joseph |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Drax: Speech Recognition with Discrete Flow Matching
por: Navon, Aviv, et al.
Publicado: (2025)
por: Navon, Aviv, et al.
Publicado: (2025)
FlowTSE: Target Speaker Extraction with Flow Matching
por: Navon, Aviv, et al.
Publicado: (2025)
por: Navon, Aviv, et al.
Publicado: (2025)
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
por: Glazer, Neta, et al.
Publicado: (2025)
por: Glazer, Neta, et al.
Publicado: (2025)
Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASR
por: Segal-Feldman, Yael, et al.
Publicado: (2024)
por: Segal-Feldman, Yael, et al.
Publicado: (2024)
Beyond Transcription: Mechanistic Interpretability in ASR
por: Glazer, Neta, et al.
Publicado: (2025)
por: Glazer, Neta, et al.
Publicado: (2025)
Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes
por: Glazer, Neta, et al.
Publicado: (2025)
por: Glazer, Neta, et al.
Publicado: (2025)
LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading
por: Yemini, Yochai, et al.
Publicado: (2023)
por: Yemini, Yochai, et al.
Publicado: (2023)
Keyword Spotting with Hyper-Matched Filters for Small Footprint Devices
por: Segal-Feldman, Yael, et al.
Publicado: (2025)
por: Segal-Feldman, Yael, et al.
Publicado: (2025)
Test-Time Adaptation for Speech Emotion Recognition
por: Dong, Jiaheng, et al.
Publicado: (2026)
por: Dong, Jiaheng, et al.
Publicado: (2026)
Adapting Automatic Speech Recognition for Accented Air Traffic Control Communications
por: Wee, Marcus Yu Zhe, et al.
Publicado: (2025)
por: Wee, Marcus Yu Zhe, et al.
Publicado: (2025)
Parameter Efficient Finetuning for Speech Emotion Recognition and Domain Adaptation
por: Lashkarashvili, Nineli, et al.
Publicado: (2024)
por: Lashkarashvili, Nineli, et al.
Publicado: (2024)
Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models
por: Feng, Chen, et al.
Publicado: (2025)
por: Feng, Chen, et al.
Publicado: (2025)
Gradient Norm-based Fine-Tuning for Backdoor Defense in Automatic Speech Recognition
por: Zhou, Nanjun, et al.
Publicado: (2025)
por: Zhou, Nanjun, et al.
Publicado: (2025)
Enhanced ASR Robustness to Packet Loss with a Front-End Adaptation Network
por: Dissen, Yehoshua, et al.
Publicado: (2024)
por: Dissen, Yehoshua, et al.
Publicado: (2024)
Transcription-Free Fine-Tuning of Speech Separation Models for Noisy and Reverberant Multi-Speaker Automatic Speech Recognition
por: Ravenscroft, William, et al.
Publicado: (2024)
por: Ravenscroft, William, et al.
Publicado: (2024)
WhisperNER: Unified Open Named Entity and Speech Recognition
por: Ayache, Gil, et al.
Publicado: (2024)
por: Ayache, Gil, et al.
Publicado: (2024)
Improving Speaker-independent Speech Emotion Recognition Using Dynamic Joint Distribution Adaptation
por: Lu, Cheng, et al.
Publicado: (2024)
por: Lu, Cheng, et al.
Publicado: (2024)
Impact of Speech Mode in Automatic Pathological Speech Detection
por: Sheikh, Shakeel A., et al.
Publicado: (2024)
por: Sheikh, Shakeel A., et al.
Publicado: (2024)
Sparse Binarization for Fast Keyword Spotting
por: Svirsky, Jonathan, et al.
Publicado: (2024)
por: Svirsky, Jonathan, et al.
Publicado: (2024)
WhisperRT -- Turning Whisper into a Causal Streaming Model
por: Krichli, Tomer, et al.
Publicado: (2025)
por: Krichli, Tomer, et al.
Publicado: (2025)
Multi-blank Transducers for Speech Recognition
por: Xu, Hainan, et al.
Publicado: (2022)
por: Xu, Hainan, et al.
Publicado: (2022)
Attention-Guided Adaptation for Code-Switching Speech Recognition
por: Aditya, Bobbi, et al.
Publicado: (2023)
por: Aditya, Bobbi, et al.
Publicado: (2023)
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
por: Xu, Hainan, et al.
Publicado: (2024)
por: Xu, Hainan, et al.
Publicado: (2024)
Regularizing Learnable Feature Extraction for Automatic Speech Recognition
por: Vieting, Peter, et al.
Publicado: (2025)
por: Vieting, Peter, et al.
Publicado: (2025)
Multiview Canonical Correlation Analysis for Automatic Pathological Speech Detection
por: Kaloga, Yacouba, et al.
Publicado: (2024)
por: Kaloga, Yacouba, et al.
Publicado: (2024)
Universal Robust Speech Adaptation for Cross-Domain Speech Recognition and Enhancement
por: Wang, Chien-Chun, et al.
Publicado: (2026)
por: Wang, Chien-Chun, et al.
Publicado: (2026)
Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis
por: Jiang, Xilin, et al.
Publicado: (2024)
por: Jiang, Xilin, et al.
Publicado: (2024)
Adapting WavLM for Speech Emotion Recognition
por: Diatlova, Daria, et al.
Publicado: (2024)
por: Diatlova, Daria, et al.
Publicado: (2024)
On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
por: Rossenbach, Nick, et al.
Publicado: (2024)
por: Rossenbach, Nick, et al.
Publicado: (2024)
Adversarial training of Keyword Spotting to Minimize TTS Data Overfitting
por: Park, Hyun Jin, et al.
Publicado: (2024)
por: Park, Hyun Jin, et al.
Publicado: (2024)
Efficient Adapter Tuning of Pre-trained Speech Models for Automatic Speaker Verification
por: Sang, Mufan, et al.
Publicado: (2024)
por: Sang, Mufan, et al.
Publicado: (2024)
Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound
por: Tjandra, Andros, et al.
Publicado: (2025)
por: Tjandra, Andros, et al.
Publicado: (2025)
Utilizing TTS Synthesized Data for Efficient Development of Keyword Spotting Model
por: Park, Hyun Jin, et al.
Publicado: (2024)
por: Park, Hyun Jin, et al.
Publicado: (2024)
Adaptive Noise Resilient Keyword Spotting Using One-Shot Learning
por: Martinez-Rau, Luciano Sebastian, et al.
Publicado: (2025)
por: Martinez-Rau, Luciano Sebastian, et al.
Publicado: (2025)
OnDA: On-device Channel Pruning for Efficient Personalized Keyword Spotting
por: Risso, Matteo, et al.
Publicado: (2026)
por: Risso, Matteo, et al.
Publicado: (2026)
TRNet: Two-level Refinement Network leveraging Speech Enhancement for Noise Robust Speech Emotion Recognition
por: Chen, Chengxin, et al.
Publicado: (2024)
por: Chen, Chengxin, et al.
Publicado: (2024)
Enhancing Speech Emotion Recognition Through Differentiable Architecture Search
por: Rajapakshe, Thejan, et al.
Publicado: (2023)
por: Rajapakshe, Thejan, et al.
Publicado: (2023)
Enhancing Automatic Speech Recognition Through Integrated Noise Detection Architecture
por: Singh, Karamvir
Publicado: (2025)
por: Singh, Karamvir
Publicado: (2025)
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation
por: Kamahori, Keisuke, et al.
Publicado: (2025)
por: Kamahori, Keisuke, et al.
Publicado: (2025)
Examining Test-Time Adaptation for Personalized Child Speech Recognition
por: Shi, Zhonghao, et al.
Publicado: (2024)
por: Shi, Zhonghao, et al.
Publicado: (2024)
Ejemplares similares
-
Drax: Speech Recognition with Discrete Flow Matching
por: Navon, Aviv, et al.
Publicado: (2025) -
FlowTSE: Target Speaker Extraction with Flow Matching
por: Navon, Aviv, et al.
Publicado: (2025) -
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
por: Glazer, Neta, et al.
Publicado: (2025) -
Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASR
por: Segal-Feldman, Yael, et al.
Publicado: (2024) -
Beyond Transcription: Mechanistic Interpretability in ASR
por: Glazer, Neta, et al.
Publicado: (2025)