Enhancing Automatic Speech Recognition Through Integrated Noise Detection Architecture
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Singh, Karamvir |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2025)
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2025)
Enhancing Speech Emotion Recognition Through Differentiable Architecture Search
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2023)
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2023)
Automatic Speech Recognition in the Modern Era: Architectures, Training, and Evaluation
von: Nayeem, Md., et al.
Veröffentlicht: (2025)
von: Nayeem, Md., et al.
Veröffentlicht: (2025)
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
von: Akinrintoyo, Emmanuel, et al.
Veröffentlicht: (2025)
von: Akinrintoyo, Emmanuel, et al.
Veröffentlicht: (2025)
Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition
von: Singh, Satwinder, et al.
Veröffentlicht: (2025)
von: Singh, Satwinder, et al.
Veröffentlicht: (2025)
Focal Loss based Residual Convolutional Neural Network for Speech Emotion Recognition
von: Tripathi, Suraj, et al.
Veröffentlicht: (2019)
von: Tripathi, Suraj, et al.
Veröffentlicht: (2019)
Enhanced Speech Emotion Recognition with Efficient Channel Attention Guided Deep CNN-BiLSTM Framework
von: Kundu, Niloy Kumar, et al.
Veröffentlicht: (2024)
von: Kundu, Niloy Kumar, et al.
Veröffentlicht: (2024)
Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech
von: Choi, Yerin, et al.
Veröffentlicht: (2024)
von: Choi, Yerin, et al.
Veröffentlicht: (2024)
LLMs-Integrated Automatic Hate Speech Recognition Using Controllable Text Generation Models
von: Oshima, Ryutaro, et al.
Veröffentlicht: (2026)
von: Oshima, Ryutaro, et al.
Veröffentlicht: (2026)
HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
von: Ahn, Hyebin, et al.
Veröffentlicht: (2025)
von: Ahn, Hyebin, et al.
Veröffentlicht: (2025)
Towards End-to-End Training of Automatic Speech Recognition for Nigerian Pidgin
von: Rufai, Amina Mardiyyah, et al.
Veröffentlicht: (2020)
von: Rufai, Amina Mardiyyah, et al.
Veröffentlicht: (2020)
Large Language Models are Efficient Learners of Noise-Robust Speech Recognition
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Speech Command Recognition Using LogNNet Reservoir Computing for Embedded Systems
von: Izotov, Yuriy, et al.
Veröffentlicht: (2025)
von: Izotov, Yuriy, et al.
Veröffentlicht: (2025)
Speech Emotion Recognition Using CNN and Its Use Case in Digital Healthcare
von: Nigar, Nishargo
Veröffentlicht: (2024)
von: Nigar, Nishargo
Veröffentlicht: (2024)
Enhancing Speech Quality through the Integration of BGRU and Transformer Architectures
von: Alghnam, Souliman, et al.
Veröffentlicht: (2025)
von: Alghnam, Souliman, et al.
Veröffentlicht: (2025)
Task Arithmetic can Mitigate Synthetic-to-Real Gap in Automatic Speech Recognition
von: Su, Hsuan, et al.
Veröffentlicht: (2024)
von: Su, Hsuan, et al.
Veröffentlicht: (2024)
Audio-Based Pedestrian Detection in the Presence of Vehicular Noise
von: Kim, Yonghyun, et al.
Veröffentlicht: (2025)
von: Kim, Yonghyun, et al.
Veröffentlicht: (2025)
Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models
von: Zhang, Wenda, et al.
Veröffentlicht: (2026)
von: Zhang, Wenda, et al.
Veröffentlicht: (2026)
Bridging Modalities: Knowledge Distillation and Masked Training for Translating Multi-Modal Emotion Recognition to Uni-Modal, Speech-Only Emotion Recognition
von: Muaz, Muhammad, et al.
Veröffentlicht: (2024)
von: Muaz, Muhammad, et al.
Veröffentlicht: (2024)
Soft Clustering Anchors for Self-Supervised Speech Representation Learning in Joint Embedding Prediction Architectures
von: Ioannides, Georgios, et al.
Veröffentlicht: (2026)
von: Ioannides, Georgios, et al.
Veröffentlicht: (2026)
Keyword-Guided Adaptation of Automatic Speech Recognition
von: Shamsian, Aviv, et al.
Veröffentlicht: (2024)
von: Shamsian, Aviv, et al.
Veröffentlicht: (2024)
Investigating the Effectiveness of Explainability Methods in Parkinson's Detection from Speech
von: Mancini, Eleonora, et al.
Veröffentlicht: (2024)
von: Mancini, Eleonora, et al.
Veröffentlicht: (2024)
Enhancing Neural Spoken Language Recognition: An Exploration with Multilingual Datasets
von: Anidjar, Or Haim, et al.
Veröffentlicht: (2025)
von: Anidjar, Or Haim, et al.
Veröffentlicht: (2025)
DiffEditor: Enhancing Speech Editing with Semantic Enrichment and Acoustic Consistency
von: Chen, Yang, et al.
Veröffentlicht: (2024)
von: Chen, Yang, et al.
Veröffentlicht: (2024)
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
von: Shi, Hao, et al.
Veröffentlicht: (2024)
von: Shi, Hao, et al.
Veröffentlicht: (2024)
Impact of Speech Mode in Automatic Pathological Speech Detection
von: Sheikh, Shakeel A., et al.
Veröffentlicht: (2024)
von: Sheikh, Shakeel A., et al.
Veröffentlicht: (2024)
TRAMBA: A Hybrid Transformer and Mamba Architecture for Practical Audio and Bone Conduction Speech Super Resolution and Enhancement on Mobile and Wearable Platforms
von: Sui, Yueyuan, et al.
Veröffentlicht: (2024)
von: Sui, Yueyuan, et al.
Veröffentlicht: (2024)
EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition
von: Wu, Linzhi, et al.
Veröffentlicht: (2026)
von: Wu, Linzhi, et al.
Veröffentlicht: (2026)
Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance
von: Hussain, Shehzeen, et al.
Veröffentlicht: (2025)
von: Hussain, Shehzeen, et al.
Veröffentlicht: (2025)
A Novel Hybrid Deep Learning Technique for Speech Emotion Detection using Feature Engineering
von: Chowdhury, Shahana Yasmin, et al.
Veröffentlicht: (2025)
von: Chowdhury, Shahana Yasmin, et al.
Veröffentlicht: (2025)
Do we really need Self-Attention for Streaming Automatic Speech Recognition?
von: Dkhissi, Youness, et al.
Veröffentlicht: (2026)
von: Dkhissi, Youness, et al.
Veröffentlicht: (2026)
Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on the Edge
von: Qin, Ruiyang, et al.
Veröffentlicht: (2024)
von: Qin, Ruiyang, et al.
Veröffentlicht: (2024)
ACES: Accent Subspaces for Coupling, Explanations, and Stress-Testing in Automatic Speech Recognition
von: Parekh, Swapnil
Veröffentlicht: (2026)
von: Parekh, Swapnil
Veröffentlicht: (2026)
Speech Unlearning
von: Cheng, Jiali, et al.
Veröffentlicht: (2025)
von: Cheng, Jiali, et al.
Veröffentlicht: (2025)
What Counts as Real? Speech Restoration and Voice Quality Conversion Pose New Challenges to Deepfake Detection
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026)
$\texttt{AVROBUSTBENCH}$: Benchmarking the Robustness of Audio-Visual Recognition Models at Test-Time
von: Maharana, Sarthak Kumar, et al.
Veröffentlicht: (2025)
von: Maharana, Sarthak Kumar, et al.
Veröffentlicht: (2025)
Detection and Forecasting of Parkinson Disease Progression from Speech Signal Features Using MultiLayer Perceptron and LSTM
von: Ali, Majid, et al.
Veröffentlicht: (2024)
von: Ali, Majid, et al.
Veröffentlicht: (2024)
Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
von: Fu, Szu-Wei, et al.
Veröffentlicht: (2024)
von: Fu, Szu-Wei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2025) -
Enhancing Speech Emotion Recognition Through Differentiable Architecture Search
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2023) -
Automatic Speech Recognition in the Modern Era: Architectures, Training, and Evaluation
von: Nayeem, Md., et al.
Veröffentlicht: (2025) -
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
von: Akinrintoyo, Emmanuel, et al.
Veröffentlicht: (2025) -
Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition
von: Singh, Satwinder, et al.
Veröffentlicht: (2025)