MT-HuBERT: Self-Supervised Mix-Training for Few-Shot Keyword Spotting in Mixed Speech
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Junming, Shi, Ying, Wang, Dong, Li, Lantian, Hamdulla, Askar |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Few-Shot Keyword Spotting from Mixed Speech
by: Yuan, Junming, et al.
Published: (2024)
by: Yuan, Junming, et al.
Published: (2024)
HuBERT-EE: Early Exiting HuBERT for Efficient Speech Recognition
by: Yoon, Ji Won, et al.
Published: (2022)
by: Yoon, Ji Won, et al.
Published: (2022)
DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective
by: Chi, Hyung Gun, et al.
Published: (2025)
by: Chi, Hyung Gun, et al.
Published: (2025)
Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction
by: Shi, Jiatong, et al.
Published: (2023)
by: Shi, Jiatong, et al.
Published: (2023)
Self-Supervised Syllable Discovery Based on Speaker-Disentangled HuBERT
by: Komatsu, Ryota, et al.
Published: (2024)
by: Komatsu, Ryota, et al.
Published: (2024)
How phonemes contribute to deep speaker models?
by: Li, Pengqi, et al.
Published: (2024)
by: Li, Pengqi, et al.
Published: (2024)
MSR-HuBERT: Self-supervised Pre-training for Adaptation to Multiple Sampling Rates
by: Huang, Zikang, et al.
Published: (2026)
by: Huang, Zikang, et al.
Published: (2026)
mHuBERT-147: A Compact Multilingual HuBERT Model
by: Boito, Marcely Zanon, et al.
Published: (2024)
by: Boito, Marcely Zanon, et al.
Published: (2024)
MelHuBERT: A simplified HuBERT on Mel spectrograms
by: Lin, Tzu-Quan, et al.
Published: (2022)
by: Lin, Tzu-Quan, et al.
Published: (2022)
Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
by: Wu, Wenxuan, et al.
Published: (2024)
by: Wu, Wenxuan, et al.
Published: (2024)
Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction
by: Chen, Xueyuan, et al.
Published: (2024)
by: Chen, Xueyuan, et al.
Published: (2024)
Phonetic and Lexical Discovery of a Canine Language using HuBERT
by: Li, Xingyuan, et al.
Published: (2024)
by: Li, Xingyuan, et al.
Published: (2024)
Speaker Emotion Recognition: Leveraging Self-Supervised Models for Feature Extraction Using Wav2Vec2 and HuBERT
by: Jafarzadeh, Pourya, et al.
Published: (2024)
by: Jafarzadeh, Pourya, et al.
Published: (2024)
Distilled HuBERT for Mobile Speech Emotion Recognition: A Cross-Corpus Validation Study
by: Ismail, Saifelden M.
Published: (2025)
by: Ismail, Saifelden M.
Published: (2025)
HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
by: Ahn, Hyebin, et al.
Published: (2025)
by: Ahn, Hyebin, et al.
Published: (2025)
Enhancing Few-shot Keyword Spotting Performance through Pre-Trained Self-supervised Speech Models
by: Gok, Alican, et al.
Published: (2025)
by: Gok, Alican, et al.
Published: (2025)
RobustSVC: HuBERT-based Melody Extractor and Adversarial Learning for Robust Singing Voice Conversion
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
EdgeSpot: Efficient and High-Performance Few-Shot Model for Keyword Spotting
by: Buyuksolak, Oguzhan, et al.
Published: (2026)
by: Buyuksolak, Oguzhan, et al.
Published: (2026)
Training-Free Multi-Step Inference for Target Speaker Extraction
by: You, Zhenghai, et al.
Published: (2026)
by: You, Zhenghai, et al.
Published: (2026)
Multi-Sample Dynamic Time Warping for Few-Shot Keyword Spotting
by: Wilkinghoff, Kevin, et al.
Published: (2024)
by: Wilkinghoff, Kevin, et al.
Published: (2024)
Iterative refinement, not training objective, makes HuBERT behave differently from wav2vec 2.0
by: Huo, Robin, et al.
Published: (2025)
by: Huo, Robin, et al.
Published: (2025)
SD-HuBERT: Sentence-Level Self-Distillation Induces Syllabic Organization in HuBERT
by: Cho, Cheol Jun, et al.
Published: (2023)
by: Cho, Cheol Jun, et al.
Published: (2023)
Quantization-Based Score Calibration for Few-Shot Keyword Spotting with Dynamic Time Warping in Noisy Environments
by: Wilkinghoff, Kevin, et al.
Published: (2025)
by: Wilkinghoff, Kevin, et al.
Published: (2025)
Keyword Mamba: Spoken Keyword Spotting with State Space Models
by: Ding, Hanyu, et al.
Published: (2025)
by: Ding, Hanyu, et al.
Published: (2025)
Serialized Output Training by Learned Dominance
by: Shi, Ying, et al.
Published: (2024)
by: Shi, Ying, et al.
Published: (2024)
Contrastive Learning With Audio Discrimination For Customizable Keyword Spotting In Continuous Speech
by: Xi, Yu, et al.
Published: (2024)
by: Xi, Yu, et al.
Published: (2024)
VisG AV-HuBERT: Viseme-Guided AV-HuBERT
by: Papadopoulos, Aristeidis, et al.
Published: (2026)
by: Papadopoulos, Aristeidis, et al.
Published: (2026)
AV-Lip-Sync+: Leveraging AV-HuBERT to Exploit Multimodal Inconsistency for Deepfake Detection of Frontal Face Videos
by: Shahzad, Sahibzada Adil, et al.
Published: (2023)
by: Shahzad, Sahibzada Adil, et al.
Published: (2023)
Disentangled Training with Adversarial Examples For Robust Small-footprint Keyword Spotting
by: Wang, Zhenyu, et al.
Published: (2024)
by: Wang, Zhenyu, et al.
Published: (2024)
Investigating the 'Autoencoder Behavior' in Speech Self-Supervised Models: a focus on HuBERT's Pretraining
by: Vielzeuf, Valentin
Published: (2024)
by: Vielzeuf, Valentin
Published: (2024)
Adaptive Noise Resilient Keyword Spotting Using One-Shot Learning
by: Martinez-Rau, Luciano Sebastian, et al.
Published: (2025)
by: Martinez-Rau, Luciano Sebastian, et al.
Published: (2025)
BabyHuBERT: Multilingual Self-Supervised Learning for Segmenting Speakers in Child-Centered Long-Form Recordings
by: Charlot, Théo, et al.
Published: (2025)
by: Charlot, Théo, et al.
Published: (2025)
Text-guided HuBERT: Self-Supervised Speech Pre-training via Generative Adversarial Networks
by: Ma, Duo, et al.
Published: (2024)
by: Ma, Duo, et al.
Published: (2024)
Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding
by: Xi, Yu, et al.
Published: (2025)
by: Xi, Yu, et al.
Published: (2025)
Contrastive Augmentation: An Unsupervised Learning Approach for Keyword Spotting in Speech Technology
by: Dai, Weinan, et al.
Published: (2024)
by: Dai, Weinan, et al.
Published: (2024)
Multichannel Keyword Spotting for Noisy Conditions
by: Saladukha, Dzmitry, et al.
Published: (2025)
by: Saladukha, Dzmitry, et al.
Published: (2025)
Effective Integration of KAN for Keyword Spotting
by: Xu, Anfeng, et al.
Published: (2024)
by: Xu, Anfeng, et al.
Published: (2024)
Robust Dual-Modal Speech Keyword Spotting for XR Headsets
by: Cai, Zhuojiang, et al.
Published: (2024)
by: Cai, Zhuojiang, et al.
Published: (2024)
Dual Data Scaling for Robust Two-Stage User-Defined Keyword Spotting
by: Ai, Zhiqi, et al.
Published: (2025)
by: Ai, Zhiqi, et al.
Published: (2025)
Does Single-channel Speech Enhancement Improve Keyword Spotting Accuracy? A Case Study
by: Brueggeman, Avamarie, et al.
Published: (2023)
by: Brueggeman, Avamarie, et al.
Published: (2023)
Similar Items
-
Few-Shot Keyword Spotting from Mixed Speech
by: Yuan, Junming, et al.
Published: (2024) -
HuBERT-EE: Early Exiting HuBERT for Efficient Speech Recognition
by: Yoon, Ji Won, et al.
Published: (2022) -
DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective
by: Chi, Hyung Gun, et al.
Published: (2025) -
Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction
by: Shi, Jiatong, et al.
Published: (2023) -
Self-Supervised Syllable Discovery Based on Speaker-Disentangled HuBERT
by: Komatsu, Ryota, et al.
Published: (2024)