From the perspective of perceptual speech quality: The robustness of frequency bands to noise
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fan, Junyi, Williamson, Donald S. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FlowDec: A flow-based full-band general audio codec with high perceptual quality
von: Welker, Simon, et al.
Veröffentlicht: (2025)
von: Welker, Simon, et al.
Veröffentlicht: (2025)
Using RLHF to align speech enhancement approaches to mean-opinion quality scores
von: Kumar, Anurag, et al.
Veröffentlicht: (2024)
von: Kumar, Anurag, et al.
Veröffentlicht: (2024)
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
von: Yan, Haoyin, et al.
Veröffentlicht: (2024)
von: Yan, Haoyin, et al.
Veröffentlicht: (2024)
Self-supervised speech representation and contextual text embedding for match-mismatch classification with EEG recording
von: Wang, Bo, et al.
Veröffentlicht: (2024)
von: Wang, Bo, et al.
Veröffentlicht: (2024)
Towards noise-robust speech inversion through multi-task learning with speech enhancement
von: Tabatabaee, Saba, et al.
Veröffentlicht: (2026)
von: Tabatabaee, Saba, et al.
Veröffentlicht: (2026)
Exploiting spatial diversity for increasing the robustness of sound source localization systems against reverberation
von: Garcia-Barrios, Guillermo, et al.
Veröffentlicht: (2024)
von: Garcia-Barrios, Guillermo, et al.
Veröffentlicht: (2024)
Speech-preserving active noise control: a deep learning approach in reverberant environments
von: Dai, Shuning
Veröffentlicht: (2026)
von: Dai, Shuning
Veröffentlicht: (2026)
Benchmarking multi-component signal processing methods in the time-frequency plane
von: Miramont, Juan M., et al.
Veröffentlicht: (2024)
von: Miramont, Juan M., et al.
Veröffentlicht: (2024)
Effects of automotive microphone frequency response characteristics and noise conditions on speech and ASR quality -- an experimental evaluation
von: Buccoli, Michele, et al.
Veröffentlicht: (2025)
von: Buccoli, Michele, et al.
Veröffentlicht: (2025)
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
von: An, Keyu, et al.
Veröffentlicht: (2024)
von: An, Keyu, et al.
Veröffentlicht: (2024)
Point Processes and spatial statistics in time-frequency analysis
von: Pascal, Barbara, et al.
Veröffentlicht: (2024)
von: Pascal, Barbara, et al.
Veröffentlicht: (2024)
Online speaker diarization of meetings guided by speech separation
von: Gruttadauria, Elio, et al.
Veröffentlicht: (2024)
von: Gruttadauria, Elio, et al.
Veröffentlicht: (2024)
SSM2Mel: State Space Model to Reconstruct Mel Spectrogram from the EEG
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
Determined blind source separation via modeling adjacent frequency band correlations in speech signals
von: Wang, Jianyu, et al.
Veröffentlicht: (2025)
von: Wang, Jianyu, et al.
Veröffentlicht: (2025)
Relationship between objective and subjective perceptual measures of speech in individuals with head and neck cancer
von: Halpern, Bence Mark, et al.
Veröffentlicht: (2025)
von: Halpern, Bence Mark, et al.
Veröffentlicht: (2025)
Array-Aware Ambisonics and HRTF Encoding for Binaural Reproduction With Wearable Arrays
von: Gayer, Yhonatan, et al.
Veröffentlicht: (2025)
von: Gayer, Yhonatan, et al.
Veröffentlicht: (2025)
AutoMashup: Automatic Music Mashups Creation
von: Delabaere, Marine, et al.
Veröffentlicht: (2025)
von: Delabaere, Marine, et al.
Veröffentlicht: (2025)
Beamforming in the Reproducing Kernel Domain Based on Spatial Differentiation
von: Iwami, Takahiro, et al.
Veröffentlicht: (2025)
von: Iwami, Takahiro, et al.
Veröffentlicht: (2025)
Tracking of Intermittent and Moving Speakers : Dataset and Metrics
von: Iatariene, Taous, et al.
Veröffentlicht: (2025)
von: Iatariene, Taous, et al.
Veröffentlicht: (2025)
Using Neurogram Similarity Index Measure (NSIM) to Model Hearing Loss and Cochlear Neural Degeneration
von: Cheema, Ahsan J., et al.
Veröffentlicht: (2025)
von: Cheema, Ahsan J., et al.
Veröffentlicht: (2025)
Aliasing-Free Neural Audio Synthesis
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
RIFT: Entropy-Optimised Fractional Wavelet Constellations for Ideal Time-Frequency Estimation
von: Cozens, James M., et al.
Veröffentlicht: (2025)
von: Cozens, James M., et al.
Veröffentlicht: (2025)
Completing Sets of Prototype Transfer Functions for Subspace-based Direction of Arrival Estimation of Multiple Speakers
von: Fejgin, Daniel, et al.
Veröffentlicht: (2025)
von: Fejgin, Daniel, et al.
Veröffentlicht: (2025)
Blind Source Separation of Radar Signals in Time Domain Using Deep Learning
von: Hinderer, Sven
Veröffentlicht: (2025)
von: Hinderer, Sven
Veröffentlicht: (2025)
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
von: Haeb-Umbach, Reinhold, et al.
Veröffentlicht: (2025)
von: Haeb-Umbach, Reinhold, et al.
Veröffentlicht: (2025)
Time-domain sound field estimation using kernel ridge regression
von: Brunnström, Jesper, et al.
Veröffentlicht: (2025)
von: Brunnström, Jesper, et al.
Veröffentlicht: (2025)
Confidence-Based Self-Training for EMG-to-Speech: Leveraging Synthetic EMG for Robust Modeling
von: Chen, Xiaodan, et al.
Veröffentlicht: (2025)
von: Chen, Xiaodan, et al.
Veröffentlicht: (2025)
DiffAU: Diffusion-Based Ambisonics Upscaling
von: Milstein, Amit, et al.
Veröffentlicht: (2025)
von: Milstein, Amit, et al.
Veröffentlicht: (2025)
Optimal Scalogram for Computational Complexity Reduction in Acoustic Recognition Using Deep Learning
von: Phan, Dang Thoai, et al.
Veröffentlicht: (2025)
von: Phan, Dang Thoai, et al.
Veröffentlicht: (2025)
Reverberation-based Features for Sound Event Localization and Detection with Distance Estimation
von: Berghi, Davide, et al.
Veröffentlicht: (2025)
von: Berghi, Davide, et al.
Veröffentlicht: (2025)
Non-locally averaged pruned reassigned spectrograms: a tool for glottal pulse visualization and analysis
von: Griswold, Gabriel J., et al.
Veröffentlicht: (2025)
von: Griswold, Gabriel J., et al.
Veröffentlicht: (2025)
STNet: Prediction of Underwater Sound Speed Profiles with An Advanced Semi-Transformer Neural Network
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
Large Language Model-based Nonnegative Matrix Factorization For Cardiorespiratory Sound Separation
von: Torabi, Yasaman, et al.
Veröffentlicht: (2025)
von: Torabi, Yasaman, et al.
Veröffentlicht: (2025)
Note-Level Singing Melody Transcription for Time-Aligned Musical Score Generation
von: Kim, Leekyung, et al.
Veröffentlicht: (2025)
von: Kim, Leekyung, et al.
Veröffentlicht: (2025)
Audio signal interpolation using optimal transportation of spectrograms
von: Valdivia, David, et al.
Veröffentlicht: (2025)
von: Valdivia, David, et al.
Veröffentlicht: (2025)
30+ Years of Source Separation Research: Achievements and Future Challenges
von: Araki, Shoko, et al.
Veröffentlicht: (2025)
von: Araki, Shoko, et al.
Veröffentlicht: (2025)
BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM
von: Gong, Xun, et al.
Veröffentlicht: (2025)
von: Gong, Xun, et al.
Veröffentlicht: (2025)
Physics-Informed Direction-Aware Neural Acoustic Fields
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
Low-Complexity Neural Wind Noise Reduction for Audio Recordings
von: Eftekhari, Hesam, et al.
Veröffentlicht: (2025)
von: Eftekhari, Hesam, et al.
Veröffentlicht: (2025)
Audio Compression using Periodic Gabor with Biorthogonal Exchange: Implementation Using the Zak Transform
von: Alimi, Roger, et al.
Veröffentlicht: (2025)
von: Alimi, Roger, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FlowDec: A flow-based full-band general audio codec with high perceptual quality
von: Welker, Simon, et al.
Veröffentlicht: (2025) -
Using RLHF to align speech enhancement approaches to mean-opinion quality scores
von: Kumar, Anurag, et al.
Veröffentlicht: (2024) -
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
von: Yan, Haoyin, et al.
Veröffentlicht: (2024) -
Self-supervised speech representation and contextual text embedding for match-mismatch classification with EEG recording
von: Wang, Bo, et al.
Veröffentlicht: (2024) -
Towards noise-robust speech inversion through multi-task learning with speech enhancement
von: Tabatabaee, Saba, et al.
Veröffentlicht: (2026)