Improving vision-inspired keyword spotting using dynamic module skipping in streaming conformer encoder
Fuente:
arXiv
Saved in:
| Main Authors: | Bittar, Alexandre, Dixon, Paul, Samragh, Mohammad, Nishu, Kumari, Naik, Devang |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SLiCK: Exploiting Subsequences for Length-Constrained Keyword Spotting
by: Nishu, Kumari, et al.
Published: (2024)
by: Nishu, Kumari, et al.
Published: (2024)
Toward noise-robust whisper keyword spotting on headphones with in-earcup microphone and curriculum learning
by: Yang, Qiaoyu
Published: (2025)
by: Yang, Qiaoyu
Published: (2025)
Boosting keyword spotting through on-device learnable user speech characteristics
by: Cioflan, Cristian, et al.
Published: (2024)
by: Cioflan, Cristian, et al.
Published: (2024)
The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language
by: Zhu, Jian, et al.
Published: (2023)
by: Zhu, Jian, et al.
Published: (2023)
Open vocabulary keyword spotting through transfer learning from speech synthesis
by: V, Kesavaraj, et al.
Published: (2024)
by: V, Kesavaraj, et al.
Published: (2024)
Reconstructing the Charlie Parker Omnibook using an audio-to-score automatic transcription pipeline
by: Riley, Xavier, et al.
Published: (2024)
by: Riley, Xavier, et al.
Published: (2024)
Enhanced Automatic Drum Transcription via Drum Stem Source Separation
by: Riley, Xavier, et al.
Published: (2025)
by: Riley, Xavier, et al.
Published: (2025)
Single-channel speech enhancement by using psychoacoustical model inspired fusion framework
by: Samui, Suman
Published: (2022)
by: Samui, Suman
Published: (2022)
Blind estimation of audio effects using an auto-encoder approach and differentiable digital signal processing
by: Peladeau, Côme, et al.
Published: (2023)
by: Peladeau, Côme, et al.
Published: (2023)
Hardware-accelerated graph neural networks: an alternative approach for neuromorphic event-based audio classification and keyword spotting on SoC FPGA
by: Jeziorek, Kamil, et al.
Published: (2026)
by: Jeziorek, Kamil, et al.
Published: (2026)
Neural Ambisonics encoding for compact irregular microphone arrays
by: Heikkinen, Mikko, et al.
Published: (2024)
by: Heikkinen, Mikko, et al.
Published: (2024)
Distilling a speech and music encoder with task arithmetic
by: Ritter-Gutierrez, Fabian, et al.
Published: (2025)
by: Ritter-Gutierrez, Fabian, et al.
Published: (2025)
Scaling up masked audio encoder learning for general audio classification
by: Dinkel, Heinrich, et al.
Published: (2024)
by: Dinkel, Heinrich, et al.
Published: (2024)
Spoken language change detection inspired by speaker change detection
by: Mishra, Jagabandhu, et al.
Published: (2023)
by: Mishra, Jagabandhu, et al.
Published: (2023)
GAPS: A Large and Diverse Classical Guitar Dataset and Benchmark Transcription Model
by: Riley, Xavier, et al.
Published: (2024)
by: Riley, Xavier, et al.
Published: (2024)
Improving Music Source Separation with Diffusion and Consistency Refinement
by: Karchkhadze, Tornike, et al.
Published: (2024)
by: Karchkhadze, Tornike, et al.
Published: (2024)
How does the teacher rate? Observations from the NeuroPiano dataset
by: Zhang, Huan, et al.
Published: (2024)
by: Zhang, Huan, et al.
Published: (2024)
Improved direction of arrival estimations with a wearable microphone array for dynamic environments by reliability weighting
by: Mitchell, Daniel A., et al.
Published: (2024)
by: Mitchell, Daniel A., et al.
Published: (2024)
DualSep: A Light-weight dual-encoder convolutional recurrent network for real-time in-car speech separation
by: Wang, Ziqian, et al.
Published: (2024)
by: Wang, Ziqian, et al.
Published: (2024)
Sequential Editing for Lifelong Training of Speech Recognition Models
by: Kulshreshtha, Devang, et al.
Published: (2024)
by: Kulshreshtha, Devang, et al.
Published: (2024)
AudioRepInceptionNeXt: A lightweight single-stream architecture for efficient audio recognition
by: Lau, Kin Wai, et al.
Published: (2024)
by: Lau, Kin Wai, et al.
Published: (2024)
Resnet-conformer network with shared weights and attention mechanism for sound event localization, detection, and distance estimation
by: Vo, Quoc Thinh, et al.
Published: (2025)
by: Vo, Quoc Thinh, et al.
Published: (2025)
Improving Audio Spectrogram Transformers for Sound Event Detection Through Multi-Stage Training
by: Schmid, Florian, et al.
Published: (2024)
by: Schmid, Florian, et al.
Published: (2024)
Improving Whispered Speech Recognition Performance using Pseudo-whispered based Data Augmentation
by: Lin, Zhaofeng, et al.
Published: (2023)
by: Lin, Zhaofeng, et al.
Published: (2023)
Revisiting and Improving Scoring Fusion for Spoofing-aware Speaker Verification Using Compositional Data Analysis
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Moonbeam: A MIDI Foundation Model Using Both Absolute and Relative Music Attributes
by: Guo, Zixun, et al.
Published: (2025)
by: Guo, Zixun, et al.
Published: (2025)
Music to Dance as Language Translation using Sequence Models
by: Correia, André, et al.
Published: (2024)
by: Correia, André, et al.
Published: (2024)
A new XML conversion process for mensural music encoding : CMME\_to\_MEI (via Verovio)
by: Fiala, David, et al.
Published: (2025)
by: Fiala, David, et al.
Published: (2025)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
by: Ducorroy, Alexandre, et al.
Published: (2025)
by: Ducorroy, Alexandre, et al.
Published: (2025)
EmoFormer: A Text-Independent Speech Emotion Recognition using a Hybrid Transformer-CNN model
by: Hasan, Rashedul, et al.
Published: (2025)
by: Hasan, Rashedul, et al.
Published: (2025)
In-context learning capabilities of Large Language Models to detect suicide risk among adolescents from speech transcripts
by: Roquefort, Filomene, et al.
Published: (2025)
by: Roquefort, Filomene, et al.
Published: (2025)
Cacophony: An Improved Contrastive Audio-Text Model
by: Zhu, Ge, et al.
Published: (2024)
by: Zhu, Ge, et al.
Published: (2024)
Improving Audio Question Answering with Variational Inference
by: Chen, Haolin
Published: (2026)
by: Chen, Haolin
Published: (2026)
Improving Neural Pitch Estimation with SWIPE Kernels
by: Marttila, David, et al.
Published: (2025)
by: Marttila, David, et al.
Published: (2025)
Unsupervised Improved MVDR Beamforming for Sound Enhancement
by: Kealey, Jacob, et al.
Published: (2024)
by: Kealey, Jacob, et al.
Published: (2024)
Phonetic Richness for Improved Automatic Speaker Verification
by: Klein, Nicholas, et al.
Published: (2024)
by: Klein, Nicholas, et al.
Published: (2024)
High Resolution Guitar Transcription via Domain Adaptation
by: Riley, Xavier, et al.
Published: (2024)
by: Riley, Xavier, et al.
Published: (2024)
Multitaper mel-spectrograms for keyword spotting
by: de Souza, Douglas Baptista, et al.
Published: (2024)
by: de Souza, Douglas Baptista, et al.
Published: (2024)
Zipformer: A faster and better encoder for automatic speech recognition
by: Yao, Zengwei, et al.
Published: (2023)
by: Yao, Zengwei, et al.
Published: (2023)
Testing chatbots on the creation of encoders for audio conditioned image generation
by: León, Jorge E., et al.
Published: (2025)
by: León, Jorge E., et al.
Published: (2025)
Similar Items
-
SLiCK: Exploiting Subsequences for Length-Constrained Keyword Spotting
by: Nishu, Kumari, et al.
Published: (2024) -
Toward noise-robust whisper keyword spotting on headphones with in-earcup microphone and curriculum learning
by: Yang, Qiaoyu
Published: (2025) -
Boosting keyword spotting through on-device learnable user speech characteristics
by: Cioflan, Cristian, et al.
Published: (2024) -
The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language
by: Zhu, Jian, et al.
Published: (2023) -
Open vocabulary keyword spotting through transfer learning from speech synthesis
by: V, Kesavaraj, et al.
Published: (2024)