RepCNN: Micro-sized, Mighty Models for Wakeword Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Kundu, Arnav, Nayak, Prateeth, Padmanabhan, Priyanka, Naik, Devang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dhvani: A Weakly-supervised Phonemic Error Detection and Personalized Feedback System for Hindi
by: Rustagi, Arnav, et al.
Published: (2025)
by: Rustagi, Arnav, et al.
Published: (2025)
Enhanced Speech Emotion Recognition with Efficient Channel Attention Guided Deep CNN-BiLSTM Framework
by: Kundu, Niloy Kumar, et al.
Published: (2024)
by: Kundu, Niloy Kumar, et al.
Published: (2024)
An End-to-End Approach for Korean Wakeword Systems with Speaker Authentication
by: Seo, Geonwoo
Published: (2025)
by: Seo, Geonwoo
Published: (2025)
Explainability of CNN Based Classification Models for Acoustic Signal
by: Faruqui, Zubair, et al.
Published: (2025)
by: Faruqui, Zubair, et al.
Published: (2025)
SLiCK: Exploiting Subsequences for Length-Constrained Keyword Spotting
by: Nishu, Kumari, et al.
Published: (2024)
by: Nishu, Kumari, et al.
Published: (2024)
Controllable Prosody Generation With Partial Inputs
by: Iliescu, Dan Andrei, et al.
Published: (2023)
by: Iliescu, Dan Andrei, et al.
Published: (2023)
Speech Emotion Recognition Using CNN and Its Use Case in Digital Healthcare
by: Nigar, Nishargo
Published: (2024)
by: Nigar, Nishargo
Published: (2024)
Comparative Analysis of CNN and Transformer Architectures with Heart Cycle Normalization for Automated Phonocardiogram Classification
by: Sondermann, Martin, et al.
Published: (2025)
by: Sondermann, Martin, et al.
Published: (2025)
Vocal Tract Length Warped Features for Spoken Keyword Spotting
by: Sarkar, Achintya kr., et al.
Published: (2025)
by: Sarkar, Achintya kr., et al.
Published: (2025)
PPINtonus: Early Detection of Parkinson's Disease Using Deep-Learning Tonal Analysis
by: Reddy, Varun
Published: (2024)
by: Reddy, Varun
Published: (2024)
MixRep: Hidden Representation Mixup for Low-Resource Speech Recognition
by: Xie, Jiamin, et al.
Published: (2023)
by: Xie, Jiamin, et al.
Published: (2023)
A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection
by: Ali, Hashim, et al.
Published: (2026)
by: Ali, Hashim, et al.
Published: (2026)
Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models
by: Huang, Wei-Ping, et al.
Published: (2026)
by: Huang, Wei-Ping, et al.
Published: (2026)
Efficient Parallel Audio Generation using Group Masked Language Modeling
by: Jeong, Myeonghun, et al.
Published: (2024)
by: Jeong, Myeonghun, et al.
Published: (2024)
DeepEmoNet: Building Machine Learning Models for Automatic Emotion Recognition in Human Speeches
by: Vu, Tai
Published: (2025)
by: Vu, Tai
Published: (2025)
GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model
by: Li, Haoyang, et al.
Published: (2025)
by: Li, Haoyang, et al.
Published: (2025)
AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification
by: Siddiqui, Md. Saiful Bari, et al.
Published: (2025)
by: Siddiqui, Md. Saiful Bari, et al.
Published: (2025)
Advancing Marine Bioacoustics with Deep Generative Models: A Hybrid Augmentation Strategy for Southern Resident Killer Whale Detection
by: Padovese, Bruno, et al.
Published: (2025)
by: Padovese, Bruno, et al.
Published: (2025)
Audio-Based Pedestrian Detection in the Presence of Vehicular Noise
by: Kim, Yonghyun, et al.
Published: (2025)
by: Kim, Yonghyun, et al.
Published: (2025)
Investigating the Effectiveness of Explainability Methods in Parkinson's Detection from Speech
by: Mancini, Eleonora, et al.
Published: (2024)
by: Mancini, Eleonora, et al.
Published: (2024)
SwiftF0: Fast and Accurate Monophonic Pitch Detection
by: Nieradzik, Lars
Published: (2025)
by: Nieradzik, Lars
Published: (2025)
Evaluating Fake Music Detection Performance Under Audio Augmentations
by: Sroka, Tomasz, et al.
Published: (2025)
by: Sroka, Tomasz, et al.
Published: (2025)
Enhancing Automatic Speech Recognition Through Integrated Noise Detection Architecture
by: Singh, Karamvir
Published: (2025)
by: Singh, Karamvir
Published: (2025)
HyWA: Hypernetwork Weight Adapting Personalized Voice Activity Detection
by: Nejad, Mahsa Ghazvini, et al.
Published: (2025)
by: Nejad, Mahsa Ghazvini, et al.
Published: (2025)
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
by: Akinrintoyo, Emmanuel, et al.
Published: (2025)
by: Akinrintoyo, Emmanuel, et al.
Published: (2025)
Hybrid Disagreement-Diversity Active Learning for Bioacoustic Sound Event Detection
by: Zhang, Shiqi, et al.
Published: (2025)
by: Zhang, Shiqi, et al.
Published: (2025)
Music Plagiarism Detection: Problem Formulation and a Segment-based Solution
by: Go, Seonghyeon, et al.
Published: (2026)
by: Go, Seonghyeon, et al.
Published: (2026)
DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention Detection
by: Yan, Sheng, et al.
Published: (2024)
by: Yan, Sheng, et al.
Published: (2024)
Lightweight Hopfield Neural Networks for Bioacoustic Detection and Call Monitoring of Captive Primates
by: Lomas, Wendy, et al.
Published: (2025)
by: Lomas, Wendy, et al.
Published: (2025)
Towards Human-in-the-Loop Onset Detection: A Transfer Learning Approach for Maracatu
by: Pinto, António Sá
Published: (2025)
by: Pinto, António Sá
Published: (2025)
Deepfake Detection of Singing Voices With Whisper Encodings
by: Sharma, Falguni, et al.
Published: (2025)
by: Sharma, Falguni, et al.
Published: (2025)
Reproducible Machine Learning-based Voice Pathology Detection: Introducing the Pitch Difference Feature
by: Vrba, Jan, et al.
Published: (2024)
by: Vrba, Jan, et al.
Published: (2024)
Weakly Supervised Detection and Temporal Localization of Whale Calls in Long-Duration Bioacoustic Data
by: Nihal, Ragib Amin, et al.
Published: (2025)
by: Nihal, Ragib Amin, et al.
Published: (2025)
MADUV: The 1st INTERSPEECH Mice Autism Detection via Ultrasound Vocalization Challenge
by: Yang, Zijiang, et al.
Published: (2025)
by: Yang, Zijiang, et al.
Published: (2025)
MIMII-Agent: Leveraging LLMs with Function Calling for Relative Evaluation of Anomalous Sound Detection
by: Purohit, Harsh, et al.
Published: (2025)
by: Purohit, Harsh, et al.
Published: (2025)
A Novel Hybrid Deep Learning Technique for Speech Emotion Detection using Feature Engineering
by: Chowdhury, Shahana Yasmin, et al.
Published: (2025)
by: Chowdhury, Shahana Yasmin, et al.
Published: (2025)
Lightweight Joint Audio-Visual Deepfake Detection via Single-Stream Multi-Modal Learning Framework
by: Zhang, Kuiyuan, et al.
Published: (2025)
by: Zhang, Kuiyuan, et al.
Published: (2025)
What Counts as Real? Speech Restoration and Voice Quality Conversion Pose New Challenges to Deepfake Detection
by: Satish, Shree Harsha Bokkahalli, et al.
Published: (2026)
by: Satish, Shree Harsha Bokkahalli, et al.
Published: (2026)
Cluster and Separate: a GNN Approach to Voice and Staff Prediction for Score Engraving
by: Foscarin, Francesco, et al.
Published: (2024)
by: Foscarin, Francesco, et al.
Published: (2024)
Multi-Microphone Speech Emotion Recognition using the Hierarchical Token-semantic Audio Transformer Architecture
by: Cohen, Ohad, et al.
Published: (2024)
by: Cohen, Ohad, et al.
Published: (2024)
Similar Items
-
Dhvani: A Weakly-supervised Phonemic Error Detection and Personalized Feedback System for Hindi
by: Rustagi, Arnav, et al.
Published: (2025) -
Enhanced Speech Emotion Recognition with Efficient Channel Attention Guided Deep CNN-BiLSTM Framework
by: Kundu, Niloy Kumar, et al.
Published: (2024) -
An End-to-End Approach for Korean Wakeword Systems with Speaker Authentication
by: Seo, Geonwoo
Published: (2025) -
Explainability of CNN Based Classification Models for Acoustic Signal
by: Faruqui, Zubair, et al.
Published: (2025) -
SLiCK: Exploiting Subsequences for Length-Constrained Keyword Spotting
by: Nishu, Kumari, et al.
Published: (2024)