Low-resource keyword spotting using contrastively trained transformer acoustic word embeddings
Fuente:
arXiv
Saved in:
| Main Authors: | Herreilers, Julian, Jacobs, Christiaan, Niesler, Thomas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multilingual acoustic word embeddings for zero-resource languages
by: Jacobs, Christiaan
Published: (2024)
by: Jacobs, Christiaan
Published: (2024)
Multitaper mel-spectrograms for keyword spotting
by: de Souza, Douglas Baptista, et al.
Published: (2024)
by: de Souza, Douglas Baptista, et al.
Published: (2024)
Learning to rumble: Automated elephant call classification, detection and endpointing using deep architectures
by: Geldenhuys, Christiaan M., et al.
Published: (2024)
by: Geldenhuys, Christiaan M., et al.
Published: (2024)
From Birdsong to Rumbles: Classifying Elephant Calls with Out-of-Species Embeddings
by: Geldenhuys, Christiaan M., et al.
Published: (2026)
by: Geldenhuys, Christiaan M., et al.
Published: (2026)
Toward noise-robust whisper keyword spotting on headphones with in-earcup microphone and curriculum learning
by: Yang, Qiaoyu
Published: (2025)
by: Yang, Qiaoyu
Published: (2025)
Boosting keyword spotting through on-device learnable user speech characteristics
by: Cioflan, Cristian, et al.
Published: (2024)
by: Cioflan, Cristian, et al.
Published: (2024)
The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language
by: Zhu, Jian, et al.
Published: (2023)
by: Zhu, Jian, et al.
Published: (2023)
Open vocabulary keyword spotting through transfer learning from speech synthesis
by: V, Kesavaraj, et al.
Published: (2024)
by: V, Kesavaraj, et al.
Published: (2024)
WhaleVAD-BPN: Improving Baleen Whale Call Detection with Boundary Proposal Networks and Post-processing Optimisation
by: Geldenhuys, Christiaan M., et al.
Published: (2025)
by: Geldenhuys, Christiaan M., et al.
Published: (2025)
Guiding the underwater acoustic target recognition with interpretable contrastive learning
by: Xie, Yuan, et al.
Published: (2024)
by: Xie, Yuan, et al.
Published: (2024)
Improving vision-inspired keyword spotting using dynamic module skipping in streaming conformer encoder
by: Bittar, Alexandre, et al.
Published: (2023)
by: Bittar, Alexandre, et al.
Published: (2023)
Hardware-accelerated graph neural networks: an alternative approach for neuromorphic event-based audio classification and keyword spotting on SoC FPGA
by: Jeziorek, Kamil, et al.
Published: (2026)
by: Jeziorek, Kamil, et al.
Published: (2026)
Visually grounded few-shot word learning in low-resource settings
by: Nortje, Leanne, et al.
Published: (2023)
by: Nortje, Leanne, et al.
Published: (2023)
Predicting positive transfer for improved low-resource speech recognition using acoustic pseudo-tokens
by: San, Nay, et al.
Published: (2024)
by: San, Nay, et al.
Published: (2024)
Progressive unsupervised domain adaptation for ASR using ensemble models and multi-stage training
by: Ahmad, Rehan, et al.
Published: (2024)
by: Ahmad, Rehan, et al.
Published: (2024)
Challenging margin-based speaker embedding extractors by using the variational information bottleneck
by: Stafylakis, Themos, et al.
Published: (2024)
by: Stafylakis, Themos, et al.
Published: (2024)
Complexity boosted adaptive training for better low resource ASR performance
by: Lu, Hongxuan, et al.
Published: (2024)
by: Lu, Hongxuan, et al.
Published: (2024)
From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation model
by: Lavechin, Marvin, et al.
Published: (2025)
by: Lavechin, Marvin, et al.
Published: (2025)
Complete reconstruction of the tongue contour through acoustic to articulatory inversion using real-time MRI data
by: Azzouz, Sofiane, et al.
Published: (2024)
by: Azzouz, Sofiane, et al.
Published: (2024)
Robust DOA estimation using deep acoustic imaging
by: Roman, Adrian S., et al.
Published: (2024)
by: Roman, Adrian S., et al.
Published: (2024)
Cross-lingual Data Selection Using Clip-level Acoustic Similarity for Enhancing Low-resource Automatic Speech Recognition
by: Mitsumori, Shunsuke, et al.
Published: (2025)
by: Mitsumori, Shunsuke, et al.
Published: (2025)
Whisper-SV: Adapting Whisper for Low-data-resource Speaker Verification
by: Zhang, Li, et al.
Published: (2024)
by: Zhang, Li, et al.
Published: (2024)
Investigation of perception inconsistency in speaker embedding for asynchronous voice anonymization
by: Wang, Rui, et al.
Published: (2025)
by: Wang, Rui, et al.
Published: (2025)
Automatically assessing oral narratives of Afrikaans and isiXhosa children
by: Louw, Retief, et al.
Published: (2025)
by: Louw, Retief, et al.
Published: (2025)
Geodesic interpolation of frame-wise speaker embeddings for the diarization of meeting scenarios
by: Cord-Landwehr, Tobias, et al.
Published: (2024)
by: Cord-Landwehr, Tobias, et al.
Published: (2024)
Neural acoustic multipole splatting for room impulse response synthesis
by: Baek, Geonwoo, et al.
Published: (2025)
by: Baek, Geonwoo, et al.
Published: (2025)
Evaluating pretrained speech embedding systems for dysarthria detection across heterogenous datasets
by: Wihlborg, Lovisa, et al.
Published: (2025)
by: Wihlborg, Lovisa, et al.
Published: (2025)
Cough activity detection for automatic tuberculosis screening
by: van Vüren, Joshua Jansen, et al.
Published: (2026)
by: van Vüren, Joshua Jansen, et al.
Published: (2026)
Tandem spoofing-robust automatic speaker verification based on time-domain embeddings
by: Weizman, Avishai, et al.
Published: (2024)
by: Weizman, Avishai, et al.
Published: (2024)
Improving acoustic drone detection generalization through pretraining and data augmentation
by: Reuter, Paul M., et al.
Published: (2026)
by: Reuter, Paul M., et al.
Published: (2026)
Perceptual implications of simplifying geometrical acoustics models for Ambisonics-based binaural reverberation
by: Martin, Vincent, et al.
Published: (2024)
by: Martin, Vincent, et al.
Published: (2024)
QiandaoEar22: A high quality noise dataset for identifying specific ship from multiple underwater acoustic targets using ship-radiated noise
by: Du, Xiaoyang, et al.
Published: (2024)
by: Du, Xiaoyang, et al.
Published: (2024)
ELAICHI: Enhancing Low-resource TTS by Addressing Infrequent and Low-frequency Character Bigrams
by: Anand, Srija, et al.
Published: (2024)
by: Anand, Srija, et al.
Published: (2024)
Physics-informed neural network for acoustic resonance analysis in a one-dimensional acoustic tube
by: Yokota, Kazuya, et al.
Published: (2023)
by: Yokota, Kazuya, et al.
Published: (2023)
A state-space representation of the boundary integral equation for room acoustic modelling
by: Ali, Randall, et al.
Published: (2026)
by: Ali, Randall, et al.
Published: (2026)
Spoken-Term Discovery using Discrete Speech Units
by: van Niekerk, Benjamin, et al.
Published: (2024)
by: van Niekerk, Benjamin, et al.
Published: (2024)
Post-training for Deepfake Speech Detection
by: Ge, Wanying, et al.
Published: (2025)
by: Ge, Wanying, et al.
Published: (2025)
Transcribe, Align and Segment: Creating speech datasets for low-resource languages
by: Sereda, Taras
Published: (2024)
by: Sereda, Taras
Published: (2024)
Directional reflection modeling via wavenumber-domain reflection coefficient for 3D acoustic field simulation
by: Hoshika, Satoshi, et al.
Published: (2026)
by: Hoshika, Satoshi, et al.
Published: (2026)
SelfTTS: cross-speaker style transfer through explicit embedding disentanglement and self-refinement using self-augmentation
by: Ueda, Lucas H., et al.
Published: (2026)
by: Ueda, Lucas H., et al.
Published: (2026)
Similar Items
-
Multilingual acoustic word embeddings for zero-resource languages
by: Jacobs, Christiaan
Published: (2024) -
Multitaper mel-spectrograms for keyword spotting
by: de Souza, Douglas Baptista, et al.
Published: (2024) -
Learning to rumble: Automated elephant call classification, detection and endpointing using deep architectures
by: Geldenhuys, Christiaan M., et al.
Published: (2024) -
From Birdsong to Rumbles: Classifying Elephant Calls with Out-of-Species Embeddings
by: Geldenhuys, Christiaan M., et al.
Published: (2026) -
Toward noise-robust whisper keyword spotting on headphones with in-earcup microphone and curriculum learning
by: Yang, Qiaoyu
Published: (2025)