Assessing speech quality metrics for evaluation of neural audio codecs under clean speech conditions
Fuente:
arXiv
Saved in:
| Main Authors: | Mack, Wolfgang, Topaloglu, Nezih, Lechler, Laura, Balić, Ivana, Craciun, Alexandra, Yesilbursa, Mansur, Wojcicki, Kamil |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MUSHRA-1S: A scalable and sensitive test approach for evaluating top-tier speech processing systems
by: Lechler, Laura, et al.
Published: (2025)
by: Lechler, Laura, et al.
Published: (2025)
Low-Resource Audio Codec (LRAC): 2025 Challenge Description
by: Wojcicki, Kamil, et al.
Published: (2025)
by: Wojcicki, Kamil, et al.
Published: (2025)
Modeling strategies for speech enhancement in the latent space of a neural audio codec
by: Kammoun, Sofiene, et al.
Published: (2025)
by: Kammoun, Sofiene, et al.
Published: (2025)
PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech Enhancement
by: Rong, Xiaobin, et al.
Published: (2025)
by: Rong, Xiaobin, et al.
Published: (2025)
Real-time speech enhancement in noise for throat microphone using neural audio codec as foundation model
by: Hauret, Julien, et al.
Published: (2025)
by: Hauret, Julien, et al.
Published: (2025)
Crowdsourced Multilingual Speech Intelligibility Testing
by: Lechler, Laura, et al.
Published: (2024)
by: Lechler, Laura, et al.
Published: (2024)
Crowdsourcing MUSHRA Tests in the Age of Generative Speech Technologies: A Comparative Analysis of Subjective and Objective Testing Methods
by: Lechler, Laura, et al.
Published: (2025)
by: Lechler, Laura, et al.
Published: (2025)
StuPASE: Towards Low-Hallucination Studio-Quality Generative Speech Enhancement
by: Rong, Xiaobin, et al.
Published: (2026)
by: Rong, Xiaobin, et al.
Published: (2026)
FreeCodec: A disentangled neural speech codec with fewer tokens
by: Zheng, Youqiang, et al.
Published: (2024)
by: Zheng, Youqiang, et al.
Published: (2024)
Speaker anonymization using neural audio codec language models
by: Panariello, Michele, et al.
Published: (2023)
by: Panariello, Michele, et al.
Published: (2023)
LDCodec: A high quality neural audio codec with low-complexity decoder
by: Jiang, Jiawei, et al.
Published: (2025)
by: Jiang, Jiawei, et al.
Published: (2025)
Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
by: Wu, Haibin, et al.
Published: (2024)
by: Wu, Haibin, et al.
Published: (2024)
PitchFlower: A flow-based neural audio codec with pitch controllability
by: Torres, Diego, et al.
Published: (2025)
by: Torres, Diego, et al.
Published: (2025)
EnCodecMAE: Leveraging neural codecs for universal audio representation learning
by: Pepino, Leonardo, et al.
Published: (2023)
by: Pepino, Leonardo, et al.
Published: (2023)
Good practices for evaluation of synthesized speech
by: Cooper, Erica, et al.
Published: (2025)
by: Cooper, Erica, et al.
Published: (2025)
Enhancement by postfiltering for speech and audio coding in ad-hoc sensor networks
by: Das, Sneha, et al.
Published: (2020)
by: Das, Sneha, et al.
Published: (2020)
An automatic mixing speech enhancement system for multi-track audio
by: Liu, Xiaojing, et al.
Published: (2024)
by: Liu, Xiaojing, et al.
Published: (2024)
Can large audio language models understand child stuttering speech? speech summarization, and source separation
by: Okocha, Chibuzor, et al.
Published: (2025)
by: Okocha, Chibuzor, et al.
Published: (2025)
Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data
by: Shirahata, Yuma, et al.
Published: (2024)
by: Shirahata, Yuma, et al.
Published: (2024)
Spectrogram features for audio and speech analysis
by: McLoughlin, Ian, et al.
Published: (2026)
by: McLoughlin, Ian, et al.
Published: (2026)
Acoustic characterization of speech rhythm: going beyond metrics with recurrent neural networks
by: Deloche, François, et al.
Published: (2024)
by: Deloche, François, et al.
Published: (2024)
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
by: Kwak, Doyeop, et al.
Published: (2026)
by: Kwak, Doyeop, et al.
Published: (2026)
Ensemble of classifiers for speech evaluation
by: Belokrylov, G., et al.
Published: (2024)
by: Belokrylov, G., et al.
Published: (2024)
Late fusion ensembles for speech recognition on diverse input audio representations
by: Jezidžić, Marin, et al.
Published: (2024)
by: Jezidžić, Marin, et al.
Published: (2024)
SeMaScore : a new evaluation metric for automatic speech recognition tasks
by: Sasindran, Zitha, et al.
Published: (2024)
by: Sasindran, Zitha, et al.
Published: (2024)
Lightweight speech enhancement guided target speech extraction in noisy multi-speaker scenarios
by: Huang, Ziling, et al.
Published: (2025)
by: Huang, Ziling, et al.
Published: (2025)
Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
by: Saon, George, et al.
Published: (2025)
by: Saon, George, et al.
Published: (2025)
SLM-S2ST: A multimodal language model for direct speech-to-speech translation
by: Hu, Yuxuan, et al.
Published: (2025)
by: Hu, Yuxuan, et al.
Published: (2025)
Effects of automotive microphone frequency response characteristics and noise conditions on speech and ASR quality -- an experimental evaluation
by: Buccoli, Michele, et al.
Published: (2025)
by: Buccoli, Michele, et al.
Published: (2025)
Forensic deepfake audio detection using segmental speech features
by: Yang, Tianle, et al.
Published: (2025)
by: Yang, Tianle, et al.
Published: (2025)
Spatially constrained vs. unconstrained filtering in neural spatiospectral filters for multichannel speech enhancement
by: Briegleb, Annika, et al.
Published: (2024)
by: Briegleb, Annika, et al.
Published: (2024)
Prominence-aware automatic speech recognition for conversational speech
by: Linke, Julian, et al.
Published: (2025)
by: Linke, Julian, et al.
Published: (2025)
Predicting speech intelligibility in older adults for speech enhancement using the Gammachirp Envelope Similarity Index, GESI
by: Yamamoto, Ayako, et al.
Published: (2025)
by: Yamamoto, Ayako, et al.
Published: (2025)
Towards noise-robust speech inversion through multi-task learning with speech enhancement
by: Tabatabaee, Saba, et al.
Published: (2026)
by: Tabatabaee, Saba, et al.
Published: (2026)
On the relationship between speech and hearing
by: Umesh, Srinivasan, et al.
Published: (2024)
by: Umesh, Srinivasan, et al.
Published: (2024)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
by: Ducorroy, Alexandre, et al.
Published: (2025)
by: Ducorroy, Alexandre, et al.
Published: (2025)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
by: Gowda, Harshavardhana T., et al.
Published: (2025)
by: Gowda, Harshavardhana T., et al.
Published: (2025)
Audio-visual video-to-speech synthesis with synthesized input audio
by: Kefalas, Triantafyllos, et al.
Published: (2023)
by: Kefalas, Triantafyllos, et al.
Published: (2023)
Improving child speech recognition with augmented child-like speech
by: Zhang, Yuanyuan, et al.
Published: (2024)
by: Zhang, Yuanyuan, et al.
Published: (2024)
An adaptive filter bank based neural network approach for time delay estimation and speech enhancement
by: Ma, Lu
Published: (2025)
by: Ma, Lu
Published: (2025)
Similar Items
-
MUSHRA-1S: A scalable and sensitive test approach for evaluating top-tier speech processing systems
by: Lechler, Laura, et al.
Published: (2025) -
Low-Resource Audio Codec (LRAC): 2025 Challenge Description
by: Wojcicki, Kamil, et al.
Published: (2025) -
Modeling strategies for speech enhancement in the latent space of a neural audio codec
by: Kammoun, Sofiene, et al.
Published: (2025) -
PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech Enhancement
by: Rong, Xiaobin, et al.
Published: (2025) -
Real-time speech enhancement in noise for throat microphone using neural audio codec as foundation model
by: Hauret, Julien, et al.
Published: (2025)