Multitaper mel-spectrograms for keyword spotting
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | de Souza, Douglas Baptista, Bakri, Khaled Jamal, Ferreira, Fernanda, Inacio, Juliana |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Boosting keyword spotting through on-device learnable user speech characteristics
par: Cioflan, Cristian, et autres
Publié: (2024)
par: Cioflan, Cristian, et autres
Publié: (2024)
Text-only domain adaptation for end-to-end ASR using integrated text-to-mel-spectrogram generator
par: Bataev, Vladimir, et autres
Publié: (2023)
par: Bataev, Vladimir, et autres
Publié: (2023)
On combining acoustic and modulation spectrograms in an attention LSTM-based system for speech intelligibility level classification
par: Gallardo-Antolín, Ascensión, et autres
Publié: (2024)
par: Gallardo-Antolín, Ascensión, et autres
Publié: (2024)
Comparison of spectrogram scaling in multi-label Music Genre Recognition
par: Karpiński, Bartosz, et autres
Publié: (2025)
par: Karpiński, Bartosz, et autres
Publié: (2025)
Hardware-accelerated graph neural networks: an alternative approach for neuromorphic event-based audio classification and keyword spotting on SoC FPGA
par: Jeziorek, Kamil, et autres
Publié: (2026)
par: Jeziorek, Kamil, et autres
Publié: (2026)
Low-resource keyword spotting using contrastively trained transformer acoustic word embeddings
par: Herreilers, Julian, et autres
Publié: (2025)
par: Herreilers, Julian, et autres
Publié: (2025)
Improving vision-inspired keyword spotting using dynamic module skipping in streaming conformer encoder
par: Bittar, Alexandre, et autres
Publié: (2023)
par: Bittar, Alexandre, et autres
Publié: (2023)
MelHuBERT: A simplified HuBERT on Mel spectrograms
par: Lin, Tzu-Quan, et autres
Publié: (2022)
par: Lin, Tzu-Quan, et autres
Publié: (2022)
Toward noise-robust whisper keyword spotting on headphones with in-earcup microphone and curriculum learning
par: Yang, Qiaoyu
Publié: (2025)
par: Yang, Qiaoyu
Publié: (2025)
Repurposing Image Diffusion Models for Training-Free Music Style Transfer on Mel-spectrograms
par: Wang, Heehwan, et autres
Publié: (2024)
par: Wang, Heehwan, et autres
Publié: (2024)
The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language
par: Zhu, Jian, et autres
Publié: (2023)
par: Zhu, Jian, et autres
Publié: (2023)
Creating a Good Teacher for Knowledge Distillation in Acoustic Scene Classification
par: Morocutti, Tobias, et autres
Publié: (2025)
par: Morocutti, Tobias, et autres
Publié: (2025)
Open vocabulary keyword spotting through transfer learning from speech synthesis
par: V, Kesavaraj, et autres
Publié: (2024)
par: V, Kesavaraj, et autres
Publié: (2024)
Audio signal interpolation using optimal transportation of spectrograms
par: Valdivia, David, et autres
Publié: (2025)
par: Valdivia, David, et autres
Publié: (2025)
Do we need more complex representations for structure? A comparison of note duration representation for Music Transformers
par: Souza, Gabriel, et autres
Publié: (2024)
par: Souza, Gabriel, et autres
Publié: (2024)
Device-Robust Acoustic Scene Classification via Impulse Response Augmentation
par: Morocutti, Tobias, et autres
Publié: (2023)
par: Morocutti, Tobias, et autres
Publié: (2023)
Text-Independent Speaker Identification Using Audio Looping With Margin Based Loss Functions
par: Garcia, Elliot Q C, et autres
Publié: (2025)
par: Garcia, Elliot Q C, et autres
Publié: (2025)
Who Said What? An Automated Approach to Analyzing Speech in Preschool Classrooms
par: Sun, Anchen, et autres
Publié: (2024)
par: Sun, Anchen, et autres
Publié: (2024)
AV-CrossNet: an Audiovisual Complex Spectral Mapping Network for Speech Separation By Leveraging Narrow- and Cross-Band Modeling
par: Kalkhorani, Vahid Ahmadi, et autres
Publié: (2024)
par: Kalkhorani, Vahid Ahmadi, et autres
Publié: (2024)
Wireless Earphone-based Real-Time Monitoring of Breathing Exercises: A Deep Learning Approach
par: Wazir, Hassam Khan, et autres
Publié: (2024)
par: Wazir, Hassam Khan, et autres
Publié: (2024)
Vibravox: A Dataset of French Speech Captured with Body-conduction Audio Sensors
par: Hauret, Julien, et autres
Publié: (2024)
par: Hauret, Julien, et autres
Publié: (2024)
Voice Disorder Analysis: a Transformer-based Approach
par: Koudounas, Alkis, et autres
Publié: (2024)
par: Koudounas, Alkis, et autres
Publié: (2024)
SegINR: Segment-wise Implicit Neural Representation for Sequence Alignment in Neural Text-to-Speech
par: Kim, Minchan, et autres
Publié: (2024)
par: Kim, Minchan, et autres
Publié: (2024)
The Second DISPLACE Challenge : DIarization of SPeaker and LAnguage in Conversational Environments
par: Kalluri, Shareef Babu, et autres
Publié: (2024)
par: Kalluri, Shareef Babu, et autres
Publié: (2024)
VQalAttent: a Transparent Speech Generation Pipeline based on Transformer-learned VQ-VAE Latent Space
par: Rodriguez, Armani, et autres
Publié: (2024)
par: Rodriguez, Armani, et autres
Publié: (2024)
HumekaFL: Automated Detection of Neonatal Asphyxia Using Federated Learning
par: Zantou, Pamely, et autres
Publié: (2024)
par: Zantou, Pamely, et autres
Publié: (2024)
Bayesian Parameter-Efficient Fine-Tuning for Overcoming Catastrophic Forgetting
par: Chen, Haolin, et autres
Publié: (2024)
par: Chen, Haolin, et autres
Publié: (2024)
Text-to-Speech for Unseen Speakers via Low-Complexity Discrete Unit-Based Frame Selection
par: Ulgen, Ismail Rasim, et autres
Publié: (2024)
par: Ulgen, Ismail Rasim, et autres
Publié: (2024)
SSAMBA: Self-Supervised Audio Representation Learning with Mamba State Space Model
par: Shams, Siavash, et autres
Publié: (2024)
par: Shams, Siavash, et autres
Publié: (2024)
Quartered Spectral Envelope and 1D-CNN-based Classification of Normally Phonated and Whispered Speech
par: Joysingh, S. Johanan, et autres
Publié: (2024)
par: Joysingh, S. Johanan, et autres
Publié: (2024)
MaskCycleGAN-based Whisper to Normal Speech Conversion
par: Gupta, K. Rohith, et autres
Publié: (2024)
par: Gupta, K. Rohith, et autres
Publié: (2024)
PRESENT: Zero-Shot Text-to-Prosody Control
par: Lam, Perry, et autres
Publié: (2024)
par: Lam, Perry, et autres
Publié: (2024)
Learning Multi-Target TDOA Features for Sound Event Localization and Detection
par: Berg, Axel, et autres
Publié: (2024)
par: Berg, Axel, et autres
Publié: (2024)
Rethinking Speaker Embeddings for Speech Generation: Sub-Center Modeling for Capturing Intra-Speaker Diversity
par: Ulgen, Ismail Rasim, et autres
Publié: (2024)
par: Ulgen, Ismail Rasim, et autres
Publié: (2024)
The Whole Is Bigger Than the Sum of Its Parts: Modeling Individual Annotators to Capture Emotional Variability
par: Tavernor, James, et autres
Publié: (2024)
par: Tavernor, James, et autres
Publié: (2024)
a-DCF: an architecture agnostic metric with application to spoofing-robust speaker verification
par: Shim, Hye-jin, et autres
Publié: (2024)
par: Shim, Hye-jin, et autres
Publié: (2024)
Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition
par: Lee, Hyeonseung, et autres
Publié: (2024)
par: Lee, Hyeonseung, et autres
Publié: (2024)
Exploring speech style spaces with language models: Emotional TTS without emotion labels
par: Chandra, Shreeram Suresh, et autres
Publié: (2024)
par: Chandra, Shreeram Suresh, et autres
Publié: (2024)
A Context-Based Numerical Format Prediction for a Text-To-Speech System
par: Darwesh, Yaser, et autres
Publié: (2024)
par: Darwesh, Yaser, et autres
Publié: (2024)
Literary and Colloquial Dialect Identification for Tamil using Acoustic Features
par: Nanmalar, M., et autres
Publié: (2024)
par: Nanmalar, M., et autres
Publié: (2024)
Documents similaires
-
Boosting keyword spotting through on-device learnable user speech characteristics
par: Cioflan, Cristian, et autres
Publié: (2024) -
Text-only domain adaptation for end-to-end ASR using integrated text-to-mel-spectrogram generator
par: Bataev, Vladimir, et autres
Publié: (2023) -
On combining acoustic and modulation spectrograms in an attention LSTM-based system for speech intelligibility level classification
par: Gallardo-Antolín, Ascensión, et autres
Publié: (2024) -
Comparison of spectrogram scaling in multi-label Music Genre Recognition
par: Karpiński, Bartosz, et autres
Publié: (2025) -
Hardware-accelerated graph neural networks: an alternative approach for neuromorphic event-based audio classification and keyword spotting on SoC FPGA
par: Jeziorek, Kamil, et autres
Publié: (2026)