End-to-End User-Defined Keyword Spotting using Shifted Delta Coefficients
Fuente:
arXiv
Guardado en:
| Autores principales: | V, Kesavaraj, M, Anuprabha, Vuppala, Anil Kumar |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Multi-modal Approach to Dysarthria Detection and Severity Assessment Using Speech and Text Information
por: M, Anuprabha, et al.
Publicado: (2024)
por: M, Anuprabha, et al.
Publicado: (2024)
Open vocabulary keyword spotting through transfer learning from speech synthesis
por: V, Kesavaraj, et al.
Publicado: (2024)
por: V, Kesavaraj, et al.
Publicado: (2024)
Fairness in Dysarthric Speech Synthesis: Understanding Intrinsic Bias in Dysarthric Speech Cloning using F5-TTS
por: Anuprabha, M, et al.
Publicado: (2025)
por: Anuprabha, M, et al.
Publicado: (2025)
Phoneme-Level Contrastive Learning for User-Defined Keyword Spotting with Flexible Enrollment
por: Kewei, Li, et al.
Publicado: (2024)
por: Kewei, Li, et al.
Publicado: (2024)
Disentangled Training with Adversarial Examples For Robust Small-footprint Keyword Spotting
por: Wang, Zhenyu, et al.
Publicado: (2024)
por: Wang, Zhenyu, et al.
Publicado: (2024)
Contrastive Augmentation: An Unsupervised Learning Approach for Keyword Spotting in Speech Technology
por: Dai, Weinan, et al.
Publicado: (2024)
por: Dai, Weinan, et al.
Publicado: (2024)
CTC-aligned Audio-Text Embedding for Streaming Open-vocabulary Keyword Spotting
por: Jin, Sichen, et al.
Publicado: (2024)
por: Jin, Sichen, et al.
Publicado: (2024)
From Modular to End-to-End Speaker Diarization
por: Landini, Federico
Publicado: (2024)
por: Landini, Federico
Publicado: (2024)
An End-to-End Approach for Chord-Conditioned Song Generation
por: Gao, Shuochen, et al.
Publicado: (2024)
por: Gao, Shuochen, et al.
Publicado: (2024)
End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization
por: Singh, Prachi, et al.
Publicado: (2024)
por: Singh, Prachi, et al.
Publicado: (2024)
Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection
por: Zhou, Xuanru, et al.
Publicado: (2024)
por: Zhou, Xuanru, et al.
Publicado: (2024)
ED-sKWS: Early-Decision Spiking Neural Networks for Rapid,and Energy-Efficient Keyword Spotting
por: Song, Zeyang, et al.
Publicado: (2024)
por: Song, Zeyang, et al.
Publicado: (2024)
Neuro-MSBG: An End-to-End Neural Model for Hearing Loss Simulation
por: Yuan, Hui-Guan, et al.
Publicado: (2025)
por: Yuan, Hui-Guan, et al.
Publicado: (2025)
Leveraging Speaker Embeddings in End-to-End Neural Diarization for Two-Speaker Scenarios
por: Alvarez-Trejos, Juan Ignacio, et al.
Publicado: (2024)
por: Alvarez-Trejos, Juan Ignacio, et al.
Publicado: (2024)
End-to-End Real-World Polyphonic Piano Audio-to-Score Transcription with Hierarchical Decoding
por: Zeng, Wei, et al.
Publicado: (2024)
por: Zeng, Wei, et al.
Publicado: (2024)
Wave-U-Mamba: An End-To-End Framework For High-Quality And Efficient Speech Super Resolution
por: Lee, Yongjoon, et al.
Publicado: (2024)
por: Lee, Yongjoon, et al.
Publicado: (2024)
MEBM-Phoneme: Multi-scale Enhanced BrainMagic for End-to-End MEG Phoneme Classification
por: Jinghua, Liang, et al.
Publicado: (2026)
por: Jinghua, Liang, et al.
Publicado: (2026)
Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-trained BERT
por: Dai, Dongyang, et al.
Publicado: (2025)
por: Dai, Dongyang, et al.
Publicado: (2025)
Qifusion-Net: Layer-adapted Stream/Non-stream Model for End-to-End Multi-Accent Speech Recognition
por: Chen, Jinming, et al.
Publicado: (2024)
por: Chen, Jinming, et al.
Publicado: (2024)
Vocal Tract Length Warped Features for Spoken Keyword Spotting
por: Sarkar, Achintya kr., et al.
Publicado: (2025)
por: Sarkar, Achintya kr., et al.
Publicado: (2025)
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
por: Lee, Jihwan, et al.
Publicado: (2024)
por: Lee, Jihwan, et al.
Publicado: (2024)
Keyword Mamba: Spoken Keyword Spotting with State Space Models
por: Ding, Hanyu, et al.
Publicado: (2025)
por: Ding, Hanyu, et al.
Publicado: (2025)
ProKWS: Personalized Keyword Spotting via Collaborative Learning of Phonemes and Prosody
por: Pan, Jianan, et al.
Publicado: (2026)
por: Pan, Jianan, et al.
Publicado: (2026)
Stutter-Solver: End-to-end Multi-lingual Dysfluency Detection
por: Zhou, Xuanru, et al.
Publicado: (2024)
por: Zhou, Xuanru, et al.
Publicado: (2024)
End-to-end multi-channel speaker extraction and binaural speech synthesis
por: Chi, Cheng, et al.
Publicado: (2024)
por: Chi, Cheng, et al.
Publicado: (2024)
Effective Integration of KAN for Keyword Spotting
por: Xu, Anfeng, et al.
Publicado: (2024)
por: Xu, Anfeng, et al.
Publicado: (2024)
Multichannel Keyword Spotting for Noisy Conditions
por: Saladukha, Dzmitry, et al.
Publicado: (2025)
por: Saladukha, Dzmitry, et al.
Publicado: (2025)
Assessing the Impact of Anisotropy in Neural Representations of Speech: A Case Study on Keyword Spotting
por: Wisniewski, Guillaume, et al.
Publicado: (2025)
por: Wisniewski, Guillaume, et al.
Publicado: (2025)
PCOV-KWS: Multi-task Learning for Personalized Customizable Open Vocabulary Keyword Spotting
por: Pan, Jianan, et al.
Publicado: (2026)
por: Pan, Jianan, et al.
Publicado: (2026)
Lightweight End-to-end Text-to-speech Synthesis for low resource on-device applications
por: Vecino, Biel Tura, et al.
Publicado: (2025)
por: Vecino, Biel Tura, et al.
Publicado: (2025)
Relationships between Keywords and Strong Beats in Lyrical Music
por: Liao, Callie C., et al.
Publicado: (2024)
por: Liao, Callie C., et al.
Publicado: (2024)
Recent Advances in End-to-End Simultaneous Speech Translation
por: Liu, Xiaoqian, et al.
Publicado: (2024)
por: Liu, Xiaoqian, et al.
Publicado: (2024)
Retrieval Augmented End-to-End Spoken Dialog Models
por: Wang, Mingqiu, et al.
Publicado: (2024)
por: Wang, Mingqiu, et al.
Publicado: (2024)
AdaKWS: Towards Robust Keyword Spotting with Test-Time Adaptation
por: Xiao, Yang, et al.
Publicado: (2025)
por: Xiao, Yang, et al.
Publicado: (2025)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
por: Bataev, Vladimir, et al.
Publicado: (2025)
por: Bataev, Vladimir, et al.
Publicado: (2025)
LearnAFE: Circuit-Algorithm Co-design Framework for Learnable Audio Analog Front-End
por: Hu, Jinhai, et al.
Publicado: (2025)
por: Hu, Jinhai, et al.
Publicado: (2025)
Query-by-Example Keyword Spotting Using Spectral-Temporal Graph Attentive Pooling and Multi-Task Learning
por: Wang, Zhenyu, et al.
Publicado: (2024)
por: Wang, Zhenyu, et al.
Publicado: (2024)
Effective User-defined Keyword Spotting with Dual-stage Matching, Multi-modal Enrollment, and Continual Adaptation
por: Ai, Zhiqi, et al.
Publicado: (2026)
por: Ai, Zhiqi, et al.
Publicado: (2026)
Synth4Kws: Synthesized Speech for User Defined Keyword Spotting in Low Resource Environments
por: Zhu, Pai, et al.
Publicado: (2024)
por: Zhu, Pai, et al.
Publicado: (2024)
SongPrep: A Preprocessing Framework and End-to-end Model for Full-song Structure Parsing and Lyrics Transcription
por: Tan, Wei, et al.
Publicado: (2025)
por: Tan, Wei, et al.
Publicado: (2025)
Ejemplares similares
-
A Multi-modal Approach to Dysarthria Detection and Severity Assessment Using Speech and Text Information
por: M, Anuprabha, et al.
Publicado: (2024) -
Open vocabulary keyword spotting through transfer learning from speech synthesis
por: V, Kesavaraj, et al.
Publicado: (2024) -
Fairness in Dysarthric Speech Synthesis: Understanding Intrinsic Bias in Dysarthric Speech Cloning using F5-TTS
por: Anuprabha, M, et al.
Publicado: (2025) -
Phoneme-Level Contrastive Learning for User-Defined Keyword Spotting with Flexible Enrollment
por: Kewei, Li, et al.
Publicado: (2024) -
Disentangled Training with Adversarial Examples For Robust Small-footprint Keyword Spotting
por: Wang, Zhenyu, et al.
Publicado: (2024)