MM-KWS: Multi-modal Prompts for Multilingual User-defined Keyword Spotting
Fuente:
arXiv
Saved in:
| Main Authors: | Ai, Zhiqi, Chen, Zhiyong, Xu, Shugong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Effective User-defined Keyword Spotting with Dual-stage Matching, Multi-modal Enrollment, and Continual Adaptation
by: Ai, Zhiqi, et al.
Published: (2026)
by: Ai, Zhiqi, et al.
Published: (2026)
PCOV-KWS: Multi-task Learning for Personalized Customizable Open Vocabulary Keyword Spotting
by: Pan, Jianan, et al.
Published: (2026)
by: Pan, Jianan, et al.
Published: (2026)
NTC-KWS: Noise-aware CTC for Robust Keyword Spotting
by: Xi, Yu, et al.
Published: (2024)
by: Xi, Yu, et al.
Published: (2024)
ProKWS: Personalized Keyword Spotting via Collaborative Learning of Phonemes and Prosody
by: Pan, Jianan, et al.
Published: (2026)
by: Pan, Jianan, et al.
Published: (2026)
MFA-KWS: Effective Keyword Spotting with Multi-head Frame-asynchronous Decoding
by: Xi, Yu, et al.
Published: (2025)
by: Xi, Yu, et al.
Published: (2025)
Enhancing Open-Set Speaker Identification through Rapid Tuning with Speaker Reciprocal Points and Negative Sample
by: Chen, Zhiyong, et al.
Published: (2024)
by: Chen, Zhiyong, et al.
Published: (2024)
StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis
by: Chen, Zhiyong, et al.
Published: (2024)
by: Chen, Zhiyong, et al.
Published: (2024)
TDT-KWS: Fast And Accurate Keyword Spotting Using Token-and-duration Transducer
by: Xi, Yu, et al.
Published: (2024)
by: Xi, Yu, et al.
Published: (2024)
AdaKWS: Towards Robust Keyword Spotting with Test-Time Adaptation
by: Xiao, Yang, et al.
Published: (2025)
by: Xiao, Yang, et al.
Published: (2025)
LLM-Synth4KWS: Scalable Automatic Generation and Synthesis of Confusable Data for Custom Keyword Spotting
by: Zhu, Pai, et al.
Published: (2025)
by: Zhu, Pai, et al.
Published: (2025)
EdgeSpot: Efficient and High-Performance Few-Shot Model for Keyword Spotting
by: Buyuksolak, Oguzhan, et al.
Published: (2026)
by: Buyuksolak, Oguzhan, et al.
Published: (2026)
GraphemeAug: A Systematic Approach to Synthesized Hard Negative Keyword Spotting Examples
by: Zhang, Harry, et al.
Published: (2025)
by: Zhang, Harry, et al.
Published: (2025)
Enhancing Few-shot Keyword Spotting Performance through Pre-Trained Self-supervised Speech Models
by: Gok, Alican, et al.
Published: (2025)
by: Gok, Alican, et al.
Published: (2025)
Language-Universal Speech Attributes Modeling for Zero-Shot Multilingual Spoken Keyword Recognition
by: Yen, Hao, et al.
Published: (2024)
by: Yen, Hao, et al.
Published: (2024)
ED-sKWS: Early-Decision Spiking Neural Networks for Rapid,and Energy-Efficient Keyword Spotting
by: Song, Zeyang, et al.
Published: (2024)
by: Song, Zeyang, et al.
Published: (2024)
AnalyticKWS: Towards Exemplar-Free Analytic Class Incremental Learning for Small-footprint Keyword Spotting
by: Xiao, Yang, et al.
Published: (2025)
by: Xiao, Yang, et al.
Published: (2025)
WCTC-Biasing: Retraining-free Contextual Biasing ASR with Wildcard CTC-based Keyword Spotting and Inter-layer Biasing
by: Nakagome, Yu, et al.
Published: (2025)
by: Nakagome, Yu, et al.
Published: (2025)
Effective Integration of KAN for Keyword Spotting
by: Xu, Anfeng, et al.
Published: (2024)
by: Xu, Anfeng, et al.
Published: (2024)
Multiple-Instance, Cascaded Classification for Keyword Spotting in Narrow-Band Audio
by: AbdulKader, Ahmad, et al.
Published: (2017)
by: AbdulKader, Ahmad, et al.
Published: (2017)
Phoneme-Level Contrastive Learning for User-Defined Keyword Spotting with Flexible Enrollment
by: Kewei, Li, et al.
Published: (2024)
by: Kewei, Li, et al.
Published: (2024)
Keyword Mamba: Spoken Keyword Spotting with State Space Models
by: Ding, Hanyu, et al.
Published: (2025)
by: Ding, Hanyu, et al.
Published: (2025)
Query-by-Example Keyword Spotting Using Spectral-Temporal Graph Attentive Pooling and Multi-Task Learning
by: Wang, Zhenyu, et al.
Published: (2024)
by: Wang, Zhenyu, et al.
Published: (2024)
Multichannel Keyword Spotting for Noisy Conditions
by: Saladukha, Dzmitry, et al.
Published: (2025)
by: Saladukha, Dzmitry, et al.
Published: (2025)
Assessing the Impact of Anisotropy in Neural Representations of Speech: A Case Study on Keyword Spotting
by: Wisniewski, Guillaume, et al.
Published: (2025)
by: Wisniewski, Guillaume, et al.
Published: (2025)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
by: Guan, Wenhao, et al.
Published: (2023)
by: Guan, Wenhao, et al.
Published: (2023)
Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting
by: Kashiwagi, Yosuke, et al.
Published: (2024)
by: Kashiwagi, Yosuke, et al.
Published: (2024)
Teaching a Multilingual Large Language Model to Understand Multilingual Speech via Multi-Instructional Training
by: Denisov, Pavel, et al.
Published: (2024)
by: Denisov, Pavel, et al.
Published: (2024)
Romanization Encoding For Multilingual ASR
by: Ding, Wen, et al.
Published: (2024)
by: Ding, Wen, et al.
Published: (2024)
A Literature Review of Keyword Spotting Technologies for Urdu
by: Rizvi, Syed Muhammad Aqdas
Published: (2024)
by: Rizvi, Syed Muhammad Aqdas
Published: (2024)
Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling
by: Liu, Rui, et al.
Published: (2024)
by: Liu, Rui, et al.
Published: (2024)
Streaming Keyword Spotting Boosted by Cross-layer Discrimination Consistency
by: Xi, Yu, et al.
Published: (2024)
by: Xi, Yu, et al.
Published: (2024)
HCAM -- Hierarchical Cross Attention Model for Multi-modal Emotion Recognition
by: Dutta, Soumya, et al.
Published: (2023)
by: Dutta, Soumya, et al.
Published: (2023)
A Language-Agnostic Hierarchical LoRA-MoE Architecture for CTC-based Multilingual ASR
by: Zheng, Yuang, et al.
Published: (2026)
by: Zheng, Yuang, et al.
Published: (2026)
Exploring SSL Discrete Tokens for Multilingual ASR
by: Cui, Mingyu, et al.
Published: (2024)
by: Cui, Mingyu, et al.
Published: (2024)
MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
by: Le-Duc, Khai, et al.
Published: (2024)
by: Le-Duc, Khai, et al.
Published: (2024)
M3TCM: Multi-modal Multi-task Context Model for Utterance Classification in Motivational Interviews
by: Hossain, Sayed Muddashir, et al.
Published: (2024)
by: Hossain, Sayed Muddashir, et al.
Published: (2024)
Multi-Teacher Language-Aware Knowledge Distillation for Multilingual Speech Emotion Recognition
by: Bijoy, Mehedi Hasan, et al.
Published: (2025)
by: Bijoy, Mehedi Hasan, et al.
Published: (2025)
Contrastive Learning With Audio Discrimination For Customizable Keyword Spotting In Continuous Speech
by: Xi, Yu, et al.
Published: (2024)
by: Xi, Yu, et al.
Published: (2024)
Multilingual DistilWhisper: Efficient Distillation of Multi-task Speech Models via Language-Specific Experts
by: Ferraz, Thomas Palmeira, et al.
Published: (2023)
by: Ferraz, Thomas Palmeira, et al.
Published: (2023)
Frequency & Channel Attention Network for Small Footprint Noisy Spoken Keyword Spotting
by: Lin, Yuanxi, et al.
Published: (2024)
by: Lin, Yuanxi, et al.
Published: (2024)
Similar Items
-
Effective User-defined Keyword Spotting with Dual-stage Matching, Multi-modal Enrollment, and Continual Adaptation
by: Ai, Zhiqi, et al.
Published: (2026) -
PCOV-KWS: Multi-task Learning for Personalized Customizable Open Vocabulary Keyword Spotting
by: Pan, Jianan, et al.
Published: (2026) -
NTC-KWS: Noise-aware CTC for Robust Keyword Spotting
by: Xi, Yu, et al.
Published: (2024) -
ProKWS: Personalized Keyword Spotting via Collaborative Learning of Phonemes and Prosody
by: Pan, Jianan, et al.
Published: (2026) -
MFA-KWS: Effective Keyword Spotting with Multi-head Frame-asynchronous Decoding
by: Xi, Yu, et al.
Published: (2025)