Speech Command Recognition Using LogNNet Reservoir Computing for Embedded Systems
Fuente:
arXiv
Guardado en:
| Autores principales: | Izotov, Yuriy, Velichko, Andrei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Speech Enhancement Using Continuous Embeddings of Neural Audio Codec
por: Li, Haoyang, et al.
Publicado: (2025)
por: Li, Haoyang, et al.
Publicado: (2025)
Enhancing Synthetic Training Data for Speech Commands: From ASR-Based Filtering to Domain Adaptation in SSL Latent Space
por: Quintas, Sebastião, et al.
Publicado: (2024)
por: Quintas, Sebastião, et al.
Publicado: (2024)
Speech Emotion Recognition Using CNN and Its Use Case in Digital Healthcare
por: Nigar, Nishargo
Publicado: (2024)
por: Nigar, Nishargo
Publicado: (2024)
Hello Afrika: Speech Commands in Kinyarwanda
por: Igwegbe, George, et al.
Publicado: (2025)
por: Igwegbe, George, et al.
Publicado: (2025)
Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition
por: Singh, Satwinder, et al.
Publicado: (2025)
por: Singh, Satwinder, et al.
Publicado: (2025)
Personalized Speech Enhancement Without a Separate Speaker Embedding Model
por: Pärnamaa, Tanel, et al.
Publicado: (2024)
por: Pärnamaa, Tanel, et al.
Publicado: (2024)
Enhancing Automatic Speech Recognition Through Integrated Noise Detection Architecture
por: Singh, Karamvir
Publicado: (2025)
por: Singh, Karamvir
Publicado: (2025)
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation
por: Kamahori, Keisuke, et al.
Publicado: (2025)
por: Kamahori, Keisuke, et al.
Publicado: (2025)
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
por: Akinrintoyo, Emmanuel, et al.
Publicado: (2025)
por: Akinrintoyo, Emmanuel, et al.
Publicado: (2025)
Focal Loss based Residual Convolutional Neural Network for Speech Emotion Recognition
por: Tripathi, Suraj, et al.
Publicado: (2019)
por: Tripathi, Suraj, et al.
Publicado: (2019)
Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
por: Fu, Szu-Wei, et al.
Publicado: (2024)
por: Fu, Szu-Wei, et al.
Publicado: (2024)
Imagined Speech State Classification for Robust Brain-Computer Interface
por: Ko, Byung-Kwan, et al.
Publicado: (2024)
por: Ko, Byung-Kwan, et al.
Publicado: (2024)
Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models
por: Zhang, Wenda, et al.
Publicado: (2026)
por: Zhang, Wenda, et al.
Publicado: (2026)
Bridging Modalities: Knowledge Distillation and Masked Training for Translating Multi-Modal Emotion Recognition to Uni-Modal, Speech-Only Emotion Recognition
por: Muaz, Muhammad, et al.
Publicado: (2024)
por: Muaz, Muhammad, et al.
Publicado: (2024)
Soft Clustering Anchors for Self-Supervised Speech Representation Learning in Joint Embedding Prediction Architectures
por: Ioannides, Georgios, et al.
Publicado: (2026)
por: Ioannides, Georgios, et al.
Publicado: (2026)
Enhanced Speech Emotion Recognition with Efficient Channel Attention Guided Deep CNN-BiLSTM Framework
por: Kundu, Niloy Kumar, et al.
Publicado: (2024)
por: Kundu, Niloy Kumar, et al.
Publicado: (2024)
Moonshine: Speech Recognition for Live Transcription and Voice Commands
por: Jeffries, Nat, et al.
Publicado: (2024)
por: Jeffries, Nat, et al.
Publicado: (2024)
The Unreliability of Acoustic Systems in Alzheimer's Speech Datasets with Heterogeneous Recording Conditions
por: Gauder, Lara, et al.
Publicado: (2024)
por: Gauder, Lara, et al.
Publicado: (2024)
Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition
por: Wu, Linzhi, et al.
Publicado: (2026)
por: Wu, Linzhi, et al.
Publicado: (2026)
Improving Pretrained YAMNet for Enhanced Speech Command Detection via Transfer Learning
por: Lachenani, Sidahmed, et al.
Publicado: (2025)
por: Lachenani, Sidahmed, et al.
Publicado: (2025)
Speech Unlearning
por: Cheng, Jiali, et al.
Publicado: (2025)
por: Cheng, Jiali, et al.
Publicado: (2025)
Detection and Forecasting of Parkinson Disease Progression from Speech Signal Features Using MultiLayer Perceptron and LSTM
por: Ali, Majid, et al.
Publicado: (2024)
por: Ali, Majid, et al.
Publicado: (2024)
When De-noising Hurts: A Systematic Study of Speech Enhancement Effects on Modern Medical ASR Systems
por: Chondhekar, Sujal, et al.
Publicado: (2025)
por: Chondhekar, Sujal, et al.
Publicado: (2025)
Collaborative Watermarking for Adversarial Speech Synthesis
por: Juvela, Lauri, et al.
Publicado: (2023)
por: Juvela, Lauri, et al.
Publicado: (2023)
Scaling Speech Tokenizers with Diffusion Autoencoders
por: Wang, Yuancheng, et al.
Publicado: (2026)
por: Wang, Yuancheng, et al.
Publicado: (2026)
Speaker Emotion Recognition: Leveraging Self-Supervised Models for Feature Extraction Using Wav2Vec2 and HuBERT
por: Jafarzadeh, Pourya, et al.
Publicado: (2024)
por: Jafarzadeh, Pourya, et al.
Publicado: (2024)
Explaining Deep Learning Embeddings for Speech Emotion Recognition by Predicting Interpretable Acoustic Features
por: Dixit, Satvik, et al.
Publicado: (2024)
por: Dixit, Satvik, et al.
Publicado: (2024)
Mitigating Unauthorized Speech Synthesis for Voice Protection
por: Zhang, Zhisheng, et al.
Publicado: (2024)
por: Zhang, Zhisheng, et al.
Publicado: (2024)
Schrödinger Bridge Mamba for One-Step Speech Enhancement
por: Yang, Jing, et al.
Publicado: (2025)
por: Yang, Jing, et al.
Publicado: (2025)
Multi-Metric Preference Alignment for Generative Speech Restoration
por: Zhang, Junan, et al.
Publicado: (2025)
por: Zhang, Junan, et al.
Publicado: (2025)
Emotion-Disentangled Embedding Alignment for Noise-Robust and Cross-Corpus Speech Emotion Recognition
por: Tiwari, Upasana, et al.
Publicado: (2025)
por: Tiwari, Upasana, et al.
Publicado: (2025)
CMGAN: Conformer-based Metric GAN for Speech Enhancement
por: Cao, Ruizhe, et al.
Publicado: (2022)
por: Cao, Ruizhe, et al.
Publicado: (2022)
High-Resolution Speech Restoration with Latent Diffusion Model
por: Dhyani, Tushar, et al.
Publicado: (2024)
por: Dhyani, Tushar, et al.
Publicado: (2024)
Single and Few-step Diffusion for Generative Speech Enhancement
por: Lay, Bunlong, et al.
Publicado: (2023)
por: Lay, Bunlong, et al.
Publicado: (2023)
The T05 System for The VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech
por: Baba, Kaito, et al.
Publicado: (2024)
por: Baba, Kaito, et al.
Publicado: (2024)
Leveraging Whisper Embeddings for Audio-based Lyrics Matching
por: Mancini, Eleonora, et al.
Publicado: (2025)
por: Mancini, Eleonora, et al.
Publicado: (2025)
Imperceptible Rhythm Backdoor Attacks: Exploring Rhythm Transformation for Embedding Undetectable Vulnerabilities on Speech Recognition
por: Yao, Wenhan, et al.
Publicado: (2024)
por: Yao, Wenhan, et al.
Publicado: (2024)
Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
por: Hajal, Karl El, et al.
Publicado: (2025)
por: Hajal, Karl El, et al.
Publicado: (2025)
Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR
por: Hajal, Karl El, et al.
Publicado: (2025)
por: Hajal, Karl El, et al.
Publicado: (2025)
Discrete Speech Unit Extraction via Independent Component Analysis
por: Nakamura, Tomohiko, et al.
Publicado: (2025)
por: Nakamura, Tomohiko, et al.
Publicado: (2025)
Ejemplares similares
-
Speech Enhancement Using Continuous Embeddings of Neural Audio Codec
por: Li, Haoyang, et al.
Publicado: (2025) -
Enhancing Synthetic Training Data for Speech Commands: From ASR-Based Filtering to Domain Adaptation in SSL Latent Space
por: Quintas, Sebastião, et al.
Publicado: (2024) -
Speech Emotion Recognition Using CNN and Its Use Case in Digital Healthcare
por: Nigar, Nishargo
Publicado: (2024) -
Hello Afrika: Speech Commands in Kinyarwanda
por: Igwegbe, George, et al.
Publicado: (2025) -
Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition
por: Singh, Satwinder, et al.
Publicado: (2025)