When Denoising Hinders: Revisiting Zero-Shot ASR with SAM-Audio and Whisper
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Islam, Akif, Nahar, Raufun, Hamid, Md. Ekramul |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Explainable Multi-Modal Deep Learning for Automatic Detection of Lung Diseases from Respiratory Audio Signals
von: Saky, S M Asiful Islam, et al.
Veröffentlicht: (2025)
von: Saky, S M Asiful Islam, et al.
Veröffentlicht: (2025)
Zero-Shot to Zero-Lies: Detecting Bengali Deepfake Audio through Transfer Learning
von: Samu, Most. Sharmin Sultana, et al.
Veröffentlicht: (2025)
von: Samu, Most. Sharmin Sultana, et al.
Veröffentlicht: (2025)
Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASR
von: Segal-Feldman, Yael, et al.
Veröffentlicht: (2024)
von: Segal-Feldman, Yael, et al.
Veröffentlicht: (2024)
Boosting ASR Robustness via Test-Time Reinforcement Learning with Audio-Text Semantic Rewards
von: Fang, Linghan, et al.
Veröffentlicht: (2026)
von: Fang, Linghan, et al.
Veröffentlicht: (2026)
Leveraging Whisper Embeddings for Audio-based Lyrics Matching
von: Mancini, Eleonora, et al.
Veröffentlicht: (2025)
von: Mancini, Eleonora, et al.
Veröffentlicht: (2025)
Balancing Interpretability and Performance in Motor Imagery EEG Classification: A Comparative Study of ANFIS-FBCSP-PSO and EEGNet
von: Aktar, Farjana, et al.
Veröffentlicht: (2025)
von: Aktar, Farjana, et al.
Veröffentlicht: (2025)
Whisper-RIR-Mega: A Paired Clean-Reverberant Speech Benchmark for ASR Robustness to Room Acoustics
von: Goswami, Mandip
Veröffentlicht: (2026)
von: Goswami, Mandip
Veröffentlicht: (2026)
Model Merging Improves Zero-Shot Generalization in Bioacoustic Foundation Models
von: Marincione, Davide, et al.
Veröffentlicht: (2025)
von: Marincione, Davide, et al.
Veröffentlicht: (2025)
Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio
von: Barański, Mateusz, et al.
Veröffentlicht: (2025)
von: Barański, Mateusz, et al.
Veröffentlicht: (2025)
Can Layer-wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?
von: Sinha, Abhijit, et al.
Veröffentlicht: (2025)
von: Sinha, Abhijit, et al.
Veröffentlicht: (2025)
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
von: Akinrintoyo, Emmanuel, et al.
Veröffentlicht: (2025)
von: Akinrintoyo, Emmanuel, et al.
Veröffentlicht: (2025)
Whisper-GPT: A Hybrid Representation Audio Large Language Model
von: Verma, Prateek
Veröffentlicht: (2024)
von: Verma, Prateek
Veröffentlicht: (2024)
AudioMosaic: Contrastive Masked Audio Representation Learning
von: Huang, Hanxun, et al.
Veröffentlicht: (2026)
von: Huang, Hanxun, et al.
Veröffentlicht: (2026)
TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification
von: Anand, Nishit, et al.
Veröffentlicht: (2024)
von: Anand, Nishit, et al.
Veröffentlicht: (2024)
Learning When to Think While Listening in Large Audio-Language Models
von: Song, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Song, Zhiyuan, et al.
Veröffentlicht: (2026)
AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation
von: Wang, Lu, et al.
Veröffentlicht: (2025)
von: Wang, Lu, et al.
Veröffentlicht: (2025)
Adaptability of ASR Models on Low-Resource Language: A Comparative Study of Whisper and Wav2Vec-BERT on Bangla
von: Ridoy, Md Sazzadul Islam, et al.
Veröffentlicht: (2025)
von: Ridoy, Md Sazzadul Islam, et al.
Veröffentlicht: (2025)
Geometry-Aware Optimization for Respiratory Sound Classification: Enhancing Sensitivity with SAM-Optimized Audio Spectrogram Transformers
von: Işık, Atakan, et al.
Veröffentlicht: (2025)
von: Işık, Atakan, et al.
Veröffentlicht: (2025)
Privacy-Enhancing Infant Cry Classification with Federated Transformers and Denoising Regularization
von: Owino, Geofrey, et al.
Veröffentlicht: (2025)
von: Owino, Geofrey, et al.
Veröffentlicht: (2025)
Evaluation of Deep Audio Representations for Hearables
von: Gröger, Fabian, et al.
Veröffentlicht: (2025)
von: Gröger, Fabian, et al.
Veröffentlicht: (2025)
When De-noising Hurts: A Systematic Study of Speech Enhancement Effects on Modern Medical ASR Systems
von: Chondhekar, Sujal, et al.
Veröffentlicht: (2025)
von: Chondhekar, Sujal, et al.
Veröffentlicht: (2025)
Representation-Based Data Quality Audits for Audio
von: Gonzalez-Jimenez, Alvaro, et al.
Veröffentlicht: (2025)
von: Gonzalez-Jimenez, Alvaro, et al.
Veröffentlicht: (2025)
Low-Resource Guidance for Controllable Latent Audio Diffusion
von: Novack, Zachary, et al.
Veröffentlicht: (2026)
von: Novack, Zachary, et al.
Veröffentlicht: (2026)
Exploring Token-Space Manipulation in Latent Audio Tokenizers
von: Paissan, Francesco, et al.
Veröffentlicht: (2026)
von: Paissan, Francesco, et al.
Veröffentlicht: (2026)
Structured-Noise Masked Modeling for Video, Audio and Beyond
von: Bhowmik, Aritra, et al.
Veröffentlicht: (2025)
von: Bhowmik, Aritra, et al.
Veröffentlicht: (2025)
Speech Emotion Recognition Leveraging OpenAI's Whisper Representations and Attentive Pooling Methods
von: Shendabadi, Ali, et al.
Veröffentlicht: (2026)
von: Shendabadi, Ali, et al.
Veröffentlicht: (2026)
Preference-Based Learning in Audio Applications: A Systematic Analysis
von: Broukhim, Aaron, et al.
Veröffentlicht: (2025)
von: Broukhim, Aaron, et al.
Veröffentlicht: (2025)
A Human-Inspired Decoupled Architecture for Efficient Audio Representation Learning
von: Kawano, Harunori, et al.
Veröffentlicht: (2026)
von: Kawano, Harunori, et al.
Veröffentlicht: (2026)
AUDETER: A Large-scale Dataset for Deepfake Audio Detection in Open Worlds
von: Wang, Qizhou, et al.
Veröffentlicht: (2025)
von: Wang, Qizhou, et al.
Veröffentlicht: (2025)
Phase-Aware Deep Learning with Complex-Valued CNNs for Audio Signal Applications
von: Agrawal, Naman
Veröffentlicht: (2025)
von: Agrawal, Naman
Veröffentlicht: (2025)
Real-Time Voicemail Detection in Telephony Audio Using Temporal Speech Activity Features
von: Saurav, Kumar
Veröffentlicht: (2026)
von: Saurav, Kumar
Veröffentlicht: (2026)
A$^2$-LLM: An End-to-end Conversational Audio Avatar Large Language Model
von: Hu, Xiaolin, et al.
Veröffentlicht: (2026)
von: Hu, Xiaolin, et al.
Veröffentlicht: (2026)
MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
von: Wang, Yuancheng, et al.
Veröffentlicht: (2024)
von: Wang, Yuancheng, et al.
Veröffentlicht: (2024)
DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching
von: Chen, Wei, et al.
Veröffentlicht: (2025)
von: Chen, Wei, et al.
Veröffentlicht: (2025)
QAMRO: Quality-aware Adaptive Margin Ranking Optimization for Human-aligned Assessment of Audio Generation Systems
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2025)
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2025)
kNN Retrieval for Simple and Effective Zero-Shot Multi-speaker Text-to-Speech
von: Hajal, Karl El, et al.
Veröffentlicht: (2024)
von: Hajal, Karl El, et al.
Veröffentlicht: (2024)
PARCO: Phoneme-Augmented Robust Contextual ASR via Contrastive Entity Disambiguation
von: He, Jiajun, et al.
Veröffentlicht: (2025)
von: He, Jiajun, et al.
Veröffentlicht: (2025)
Linguistic and Audio Embedding-Based Machine Learning for Alzheimer's Dementia and Mild Cognitive Impairment Detection: Insights from the PROCESS Challenge
von: Devahi, Adharsha Sam Edwin Sam, et al.
Veröffentlicht: (2025)
von: Devahi, Adharsha Sam Edwin Sam, et al.
Veröffentlicht: (2025)
Whisper-MLA: Reducing GPU Memory Consumption of ASR Models based on MHA2MLA Conversion
von: Zhang, Sen, et al.
Veröffentlicht: (2026)
von: Zhang, Sen, et al.
Veröffentlicht: (2026)
Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
von: Hajal, Karl El, et al.
Veröffentlicht: (2025)
von: Hajal, Karl El, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Explainable Multi-Modal Deep Learning for Automatic Detection of Lung Diseases from Respiratory Audio Signals
von: Saky, S M Asiful Islam, et al.
Veröffentlicht: (2025) -
Zero-Shot to Zero-Lies: Detecting Bengali Deepfake Audio through Transfer Learning
von: Samu, Most. Sharmin Sultana, et al.
Veröffentlicht: (2025) -
Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASR
von: Segal-Feldman, Yael, et al.
Veröffentlicht: (2024) -
Boosting ASR Robustness via Test-Time Reinforcement Learning with Audio-Text Semantic Rewards
von: Fang, Linghan, et al.
Veröffentlicht: (2026) -
Leveraging Whisper Embeddings for Audio-based Lyrics Matching
von: Mancini, Eleonora, et al.
Veröffentlicht: (2025)