Noise-Robust Keyword Spotting through Self-supervised Pretraining
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mørk, Jacob, Bovbjerg, Holger Severin, Kiss, Gergely, Tan, Zheng-Hua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-supervised Pretraining for Robust Personalized Voice Activity Detection in Adverse Conditions
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2023)
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2023)
Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2025)
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2025)
Learning Robust Spatial Representations from Binaural Audio through Feature Distillation
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2025)
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2025)
Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2026)
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2026)
Joint Feature and Output Distillation for Low-complexity Acoustic Scene Classification
von: Li, Haowen, et al.
Veröffentlicht: (2025)
von: Li, Haowen, et al.
Veröffentlicht: (2025)
KinSPEAK: Improving speech recognition for Kinyarwanda via semi-supervised learning methods
von: Nzeyimana, Antoine
Veröffentlicht: (2023)
von: Nzeyimana, Antoine
Veröffentlicht: (2023)
Audio-based Kinship Verification Using Age Domain Conversion
von: Sun, Qiyang, et al.
Veröffentlicht: (2024)
von: Sun, Qiyang, et al.
Veröffentlicht: (2024)
ChordSync: Conformer-Based Alignment of Chord Annotations to Music Audio
von: Poltronieri, Andrea, et al.
Veröffentlicht: (2024)
von: Poltronieri, Andrea, et al.
Veröffentlicht: (2024)
Symbolic Audio Classification via Modal Decision Tree Learning
von: Marzano, Enrico, et al.
Veröffentlicht: (2025)
von: Marzano, Enrico, et al.
Veröffentlicht: (2025)
Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
TuneGenie: Reasoning-based LLM agents for preferential music generation
von: Pandey, Amitesh, et al.
Veröffentlicht: (2025)
von: Pandey, Amitesh, et al.
Veröffentlicht: (2025)
Passive Underwater Acoustic Signal Separation based on Feature Decoupling Dual-path Network
von: Liu, Yucheng, et al.
Veröffentlicht: (2025)
von: Liu, Yucheng, et al.
Veröffentlicht: (2025)
Local Diagnostics of Continuous Normalizing Flow for Out-of-Distribution Detection
von: Cao, Xinwei, et al.
Veröffentlicht: (2026)
von: Cao, Xinwei, et al.
Veröffentlicht: (2026)
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
Experimental Study: Enhancing Voice Spoofing Detection Models with wav2vec 2.0
von: Kang, Taein, et al.
Veröffentlicht: (2024)
von: Kang, Taein, et al.
Veröffentlicht: (2024)
NTC-KWS: Noise-aware CTC for Robust Keyword Spotting
von: Xi, Yu, et al.
Veröffentlicht: (2024)
von: Xi, Yu, et al.
Veröffentlicht: (2024)
Beyond Deep Learning: Speech Segmentation and Phone Classification with Neural Assemblies
von: Adelson, Trevor, et al.
Veröffentlicht: (2026)
von: Adelson, Trevor, et al.
Veröffentlicht: (2026)
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
von: Mehta, Shivam, et al.
Veröffentlicht: (2024)
von: Mehta, Shivam, et al.
Veröffentlicht: (2024)
A Multimodal Symphony: Integrating Taste and Sound through Generative AI
von: Spanio, Matteo, et al.
Veröffentlicht: (2025)
von: Spanio, Matteo, et al.
Veröffentlicht: (2025)
Quantization-Based Score Calibration for Few-Shot Keyword Spotting with Dynamic Time Warping in Noisy Environments
von: Wilkinghoff, Kevin, et al.
Veröffentlicht: (2025)
von: Wilkinghoff, Kevin, et al.
Veröffentlicht: (2025)
Quantum-Enhanced Analysis and Grading of Vocal Performance
von: Agarwal, Rohan
Veröffentlicht: (2025)
von: Agarwal, Rohan
Veröffentlicht: (2025)
Modeling L1 Influence on L2 Pronunciation: An MFCC-Based Framework for Explainable Machine Learning and Pedagogical Feedback
von: Jahanbin, Peyman
Veröffentlicht: (2025)
von: Jahanbin, Peyman
Veröffentlicht: (2025)
CAST-TTS: A Simple Cross-Attention Framework for Unified Timbre Control in TTS
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description
von: Zheng, Zihao, et al.
Veröffentlicht: (2025)
von: Zheng, Zihao, et al.
Veröffentlicht: (2025)
FakeSound: Deepfake General Audio Detection
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
GraFPrint: A GNN-Based Approach for Audio Identification
von: Bhattacharjee, Aditya, et al.
Veröffentlicht: (2024)
von: Bhattacharjee, Aditya, et al.
Veröffentlicht: (2024)
Audio Foundation Models Outperform Symbolic Representations for Piano Performance Evaluation
von: Dhiman, Jai
Veröffentlicht: (2026)
von: Dhiman, Jai
Veröffentlicht: (2026)
Scalable Evaluation for Audio Identification via Synthetic Latent Fingerprint Generation
von: Bhattacharjee, Aditya, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Aditya, et al.
Veröffentlicht: (2025)
Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding
von: Xi, Yu, et al.
Veröffentlicht: (2025)
von: Xi, Yu, et al.
Veröffentlicht: (2025)
Prevailing Research Areas for Music AI in the Era of Foundation Models
von: Wei, Megan, et al.
Veröffentlicht: (2024)
von: Wei, Megan, et al.
Veröffentlicht: (2024)
Matcha-TTS: A fast TTS architecture with conditional flow matching
von: Mehta, Shivam, et al.
Veröffentlicht: (2023)
von: Mehta, Shivam, et al.
Veröffentlicht: (2023)
HELIX: Scaling Raw Audio Understanding with Hybrid Mamba-Attention Beyond the Quadratic Limit
von: Khushiyant, et al.
Veröffentlicht: (2026)
von: Khushiyant, et al.
Veröffentlicht: (2026)
Enhancing Few-shot Keyword Spotting Performance through Pre-Trained Self-supervised Speech Models
von: Gok, Alican, et al.
Veröffentlicht: (2025)
von: Gok, Alican, et al.
Veröffentlicht: (2025)
Fine-tuning Pre-trained Audio Models for COVID-19 Detection: A Technical Report
von: de Brito, Daniel Oliveira, et al.
Veröffentlicht: (2025)
von: de Brito, Daniel Oliveira, et al.
Veröffentlicht: (2025)
Deepfake audio as a data augmentation technique for training automatic speech to text transcription models
von: Ferreira, Alexandre R., et al.
Veröffentlicht: (2023)
von: Ferreira, Alexandre R., et al.
Veröffentlicht: (2023)
Keyword Mamba: Spoken Keyword Spotting with State Space Models
von: Ding, Hanyu, et al.
Veröffentlicht: (2025)
von: Ding, Hanyu, et al.
Veröffentlicht: (2025)
AdaKWS: Towards Robust Keyword Spotting with Test-Time Adaptation
von: Xiao, Yang, et al.
Veröffentlicht: (2025)
von: Xiao, Yang, et al.
Veröffentlicht: (2025)
STAR: Speech-to-Audio Generation via Representation Learning
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
FakeSound2: A Benchmark for Explainable and Generalizable Deepfake Sound Detection
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Self-supervised Pretraining for Robust Personalized Voice Activity Detection in Adverse Conditions
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2023) -
Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2025) -
Learning Robust Spatial Representations from Binaural Audio through Feature Distillation
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2025) -
Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2026) -
Joint Feature and Output Distillation for Low-complexity Acoustic Scene Classification
von: Li, Haowen, et al.
Veröffentlicht: (2025)