Whisper-RIR-Mega: A Paired Clean-Reverberant Speech Benchmark for ASR Robustness to Room Acoustics
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Goswami, Mandip |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RIR-Mega-Speech: A Reverberant Speech Corpus with Comprehensive Acoustic Metadata and Reproducible Evaluation
von: Goswami, Mandip
Veröffentlicht: (2026)
von: Goswami, Mandip
Veröffentlicht: (2026)
RevRIR: Joint Reverberant Speech and Room Impulse Response Embedding using Contrastive Learning with Application to Room Shape Classification
von: Bitterman, Jacob, et al.
Veröffentlicht: (2024)
von: Bitterman, Jacob, et al.
Veröffentlicht: (2024)
Acoustivision Pro: An Open-Source Interactive Platform for Room Impulse Response Analysis and Acoustic Characterization
von: Goswami, Mandip
Veröffentlicht: (2026)
von: Goswami, Mandip
Veröffentlicht: (2026)
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
von: Akinrintoyo, Emmanuel, et al.
Veröffentlicht: (2025)
von: Akinrintoyo, Emmanuel, et al.
Veröffentlicht: (2025)
Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASR
von: Segal-Feldman, Yael, et al.
Veröffentlicht: (2024)
von: Segal-Feldman, Yael, et al.
Veröffentlicht: (2024)
Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio
von: Barański, Mateusz, et al.
Veröffentlicht: (2025)
von: Barański, Mateusz, et al.
Veröffentlicht: (2025)
RIR-Mega: a large-scale simulated room impulse response dataset for machine learning and room acoustics modeling
von: Goswami, Mandip
Veröffentlicht: (2025)
von: Goswami, Mandip
Veröffentlicht: (2025)
DiffusionRIR: Room Impulse Response Interpolation using Diffusion Models
von: Della Torre, Sagi, et al.
Veröffentlicht: (2025)
von: Della Torre, Sagi, et al.
Veröffentlicht: (2025)
BeepBank-500: A Synthetic Earcon Mini-Corpus for UI Sound Research and Psychoacoustics Research
von: Goswami, Mandip
Veröffentlicht: (2025)
von: Goswami, Mandip
Veröffentlicht: (2025)
Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
von: Fu, Szu-Wei, et al.
Veröffentlicht: (2024)
von: Fu, Szu-Wei, et al.
Veröffentlicht: (2024)
Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
von: Hajal, Karl El, et al.
Veröffentlicht: (2025)
von: Hajal, Karl El, et al.
Veröffentlicht: (2025)
Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR
von: Hajal, Karl El, et al.
Veröffentlicht: (2025)
von: Hajal, Karl El, et al.
Veröffentlicht: (2025)
Automotive Sound Quality for EVs: Psychoacoustic Metrics with Reproducible AI/ML Baselines
von: Goswami, Mandip
Veröffentlicht: (2025)
von: Goswami, Mandip
Veröffentlicht: (2025)
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2025)
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2025)
CleanMel: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR
von: Shao, Nian, et al.
Veröffentlicht: (2025)
von: Shao, Nian, et al.
Veröffentlicht: (2025)
Wav2code: Restore Clean Speech Representations via Codebook Lookup for Noise-Robust ASR
von: Hu, Yuchen, et al.
Veröffentlicht: (2023)
von: Hu, Yuchen, et al.
Veröffentlicht: (2023)
EchoMark: Perceptual Acoustic Environment Transfer with Watermark-Embedded Room Impulse Response
von: Huang, Chenpei, et al.
Veröffentlicht: (2025)
von: Huang, Chenpei, et al.
Veröffentlicht: (2025)
When De-noising Hurts: A Systematic Study of Speech Enhancement Effects on Modern Medical ASR Systems
von: Chondhekar, Sujal, et al.
Veröffentlicht: (2025)
von: Chondhekar, Sujal, et al.
Veröffentlicht: (2025)
Predicting Individual Depression Symptoms from Acoustic Features During Speech
von: Rodriguez, Sebastian, et al.
Veröffentlicht: (2024)
von: Rodriguez, Sebastian, et al.
Veröffentlicht: (2024)
DiffEditor: Enhancing Speech Editing with Semantic Enrichment and Acoustic Consistency
von: Chen, Yang, et al.
Veröffentlicht: (2024)
von: Chen, Yang, et al.
Veröffentlicht: (2024)
The Unreliability of Acoustic Systems in Alzheimer's Speech Datasets with Heterogeneous Recording Conditions
von: Gauder, Lara, et al.
Veröffentlicht: (2024)
von: Gauder, Lara, et al.
Veröffentlicht: (2024)
Towards Open Respiratory Acoustic Foundation Models: Pretraining and Benchmarking
von: Zhang, Yuwei, et al.
Veröffentlicht: (2024)
von: Zhang, Yuwei, et al.
Veröffentlicht: (2024)
Leveraging Whisper Embeddings for Audio-based Lyrics Matching
von: Mancini, Eleonora, et al.
Veröffentlicht: (2025)
von: Mancini, Eleonora, et al.
Veröffentlicht: (2025)
Enhancing Synthetic Training Data for Speech Commands: From ASR-Based Filtering to Domain Adaptation in SSL Latent Space
von: Quintas, Sebastião, et al.
Veröffentlicht: (2024)
von: Quintas, Sebastião, et al.
Veröffentlicht: (2024)
Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation
von: Xie, Zhifei, et al.
Veröffentlicht: (2026)
von: Xie, Zhifei, et al.
Veröffentlicht: (2026)
Device-Robust Acoustic Scene Classification via Impulse Response Augmentation
von: Morocutti, Tobias, et al.
Veröffentlicht: (2023)
von: Morocutti, Tobias, et al.
Veröffentlicht: (2023)
VINP: Variational Bayesian Inference with Neural Speech Prior for Joint ASR-Effective Speech Dereverberation and Blind RIR Identification
von: Wang, Pengyu, et al.
Veröffentlicht: (2025)
von: Wang, Pengyu, et al.
Veröffentlicht: (2025)
Can Layer-wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?
von: Sinha, Abhijit, et al.
Veröffentlicht: (2025)
von: Sinha, Abhijit, et al.
Veröffentlicht: (2025)
Imagined Speech State Classification for Robust Brain-Computer Interface
von: Ko, Byung-Kwan, et al.
Veröffentlicht: (2024)
von: Ko, Byung-Kwan, et al.
Veröffentlicht: (2024)
Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition
von: Singh, Satwinder, et al.
Veröffentlicht: (2025)
von: Singh, Satwinder, et al.
Veröffentlicht: (2025)
Benchmarking Representations for Speech, Music, and Acoustic Events
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024)
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024)
Whisfusion: Parallel ASR Decoding via a Diffusion Transformer
von: Kwon, Taeyoun, et al.
Veröffentlicht: (2025)
von: Kwon, Taeyoun, et al.
Veröffentlicht: (2025)
Run-Time Adaptation of Neural Beamforming for Robust Speech Dereverberation and Denoising
von: Fujita, Yoto, et al.
Veröffentlicht: (2024)
von: Fujita, Yoto, et al.
Veröffentlicht: (2024)
Self-Supervised Disentangled Representation Learning for Robust Target Speech Extraction
von: Mu, Zhaoxi, et al.
Veröffentlicht: (2023)
von: Mu, Zhaoxi, et al.
Veröffentlicht: (2023)
Bridging ASR and LLMs for Dysarthric Speech Recognition: Benchmarking Self-Supervised and Generative Approaches
von: Aboeitta, Ahmed, et al.
Veröffentlicht: (2025)
von: Aboeitta, Ahmed, et al.
Veröffentlicht: (2025)
JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention
von: Ioannides, Georgios, et al.
Veröffentlicht: (2025)
von: Ioannides, Georgios, et al.
Veröffentlicht: (2025)
Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization
von: Bhuiyan, Mohammed Aman, et al.
Veröffentlicht: (2026)
von: Bhuiyan, Mohammed Aman, et al.
Veröffentlicht: (2026)
Speech Diarization and ASR with GMM
von: Sharma, Aayush Kumar, et al.
Veröffentlicht: (2023)
von: Sharma, Aayush Kumar, et al.
Veröffentlicht: (2023)
$\texttt{AVROBUSTBENCH}$: Benchmarking the Robustness of Audio-Visual Recognition Models at Test-Time
von: Maharana, Sarthak Kumar, et al.
Veröffentlicht: (2025)
von: Maharana, Sarthak Kumar, et al.
Veröffentlicht: (2025)
Quantum-Enhanced Transformers for Robust Acoustic Scene Classification in IoT Environments
von: Quan, Minh K., et al.
Veröffentlicht: (2025)
von: Quan, Minh K., et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RIR-Mega-Speech: A Reverberant Speech Corpus with Comprehensive Acoustic Metadata and Reproducible Evaluation
von: Goswami, Mandip
Veröffentlicht: (2026) -
RevRIR: Joint Reverberant Speech and Room Impulse Response Embedding using Contrastive Learning with Application to Room Shape Classification
von: Bitterman, Jacob, et al.
Veröffentlicht: (2024) -
Acoustivision Pro: An Open-Source Interactive Platform for Room Impulse Response Analysis and Acoustic Characterization
von: Goswami, Mandip
Veröffentlicht: (2026) -
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
von: Akinrintoyo, Emmanuel, et al.
Veröffentlicht: (2025) -
Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASR
von: Segal-Feldman, Yael, et al.
Veröffentlicht: (2024)