Impairments are Clustered in Latents of Deep Neural Network-based Speech Quality Models
Fuente:
arXiv
Saved in:
| Main Authors: | Cumlin, Fredrik, Liang, Xinyu, Ungureanu, Victor, Reddy, Chandan K. A., Schüldt, Christian, Chatterjee, Saikat |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multivariate Probabilistic Assessment of Speech Quality
by: Cumlin, Fredrik, et al.
Published: (2025)
by: Cumlin, Fredrik, et al.
Published: (2025)
Selection of Layers from Self-supervised Learning Models for Predicting Mean-Opinion-Score of Speech
by: Liang, Xinyu, et al.
Published: (2025)
by: Liang, Xinyu, et al.
Published: (2025)
SA-SSL-MOS: Self-supervised Learning MOS Prediction with Spectral Augmentation for Generalized Multi-Rate Speech Assessment
by: Cao, Fengyuan, et al.
Published: (2026)
by: Cao, Fengyuan, et al.
Published: (2026)
Leveraging LLMs for Scalable Non-intrusive Speech Quality Assessment
by: Cumlin, Fredrik, et al.
Published: (2025)
by: Cumlin, Fredrik, et al.
Published: (2025)
Rho-Perfect: Correlation Ceiling For Subjective Evaluation Datasets
by: Cumlin, Fredrik
Published: (2026)
by: Cumlin, Fredrik
Published: (2026)
On the Relation Between Speech Quality and Quantized Latent Representations of Neural Codecs
by: Halimeh, Mhd Modar, et al.
Published: (2025)
by: Halimeh, Mhd Modar, et al.
Published: (2025)
Investigation of Speech and Noise Latent Representations in Single-channel VAE-based Speech Enhancement
by: Li, Jiatong, et al.
Published: (2025)
by: Li, Jiatong, et al.
Published: (2025)
Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts
by: Kuhlmann, Michael, et al.
Published: (2026)
by: Kuhlmann, Michael, et al.
Published: (2026)
Enhanced Deep Speech Separation in Clustered Ad Hoc Distributed Microphone Environments
by: Kim, Jihyun, et al.
Published: (2024)
by: Kim, Jihyun, et al.
Published: (2024)
Latent Filling: Latent Space Data Augmentation for Zero-shot Speech Synthesis
by: Bae, Jae-Sung, et al.
Published: (2023)
by: Bae, Jae-Sung, et al.
Published: (2023)
A High-Quality and Low-Complexity Streamable Neural Speech Codec with Knowledge Distillation
by: Zhang, En-Wei, et al.
Published: (2025)
by: Zhang, En-Wei, et al.
Published: (2025)
Quantifying Spatial Audio Quality Impairment
by: Watcharasupat, Karn N., et al.
Published: (2023)
by: Watcharasupat, Karn N., et al.
Published: (2023)
Vision-Integrated High-Quality Neural Speech Coding
by: Guo, Yao, et al.
Published: (2025)
by: Guo, Yao, et al.
Published: (2025)
Low Bitrate High-Quality RVQGAN-based Discrete Speech Tokenizer
by: Shechtman, Slava, et al.
Published: (2024)
by: Shechtman, Slava, et al.
Published: (2024)
Learning-based A Posteriori Speech Presence Probability Estimation and Applications
by: Tao, Shuai, et al.
Published: (2025)
by: Tao, Shuai, et al.
Published: (2025)
An Efficient Neural Network for Modeling Human Auditory Neurograms for Speech
by: Zohar, Eylon, et al.
Published: (2025)
by: Zohar, Eylon, et al.
Published: (2025)
FNH-TTS: Mixture-of-Experts Duration Modeling for Robust Neural Speech Synthesis
by: Meng, Qingliang, et al.
Published: (2025)
by: Meng, Qingliang, et al.
Published: (2025)
VBx for End-to-End Neural and Clustering-based Diarization
by: Pálka, Petr, et al.
Published: (2025)
by: Pálka, Petr, et al.
Published: (2025)
Binaural Speech Enhancement Using Deep Complex Convolutional Transformer Networks
by: Tokala, Vikas, et al.
Published: (2024)
by: Tokala, Vikas, et al.
Published: (2024)
Latent Secret Spin: Keyed Orthogonal Rotations for Blind Speech Watermarking in Anisotropic Latent Spaces
by: Coletta, Emma, et al.
Published: (2026)
by: Coletta, Emma, et al.
Published: (2026)
AntiDeepFake: AI for Deep Fake Speech Recognition
by: Togootogtokh, Enkhtogtokh, et al.
Published: (2024)
by: Togootogtokh, Enkhtogtokh, et al.
Published: (2024)
Automatic Detection of Depression in Speech Using Ensemble Convolutional Neural Networks
by: Vázquez-Romero, Adrián, et al.
Published: (2024)
by: Vázquez-Romero, Adrián, et al.
Published: (2024)
Speech Quality Embeddings for Improved Detection and Classification of Degradations in Speech Signals
by: Kuhlmann, Michael, et al.
Published: (2026)
by: Kuhlmann, Michael, et al.
Published: (2026)
NeuralMultiling: A Novel Neural Architecture Search for Smartphone based Multilingual Speaker Verification
by: PN, Aravinda Reddy, et al.
Published: (2024)
by: PN, Aravinda Reddy, et al.
Published: (2024)
DAT-CFTNet: Speech Enhancement for Cochlear Implant Recipients using Attention-based Dual-Path Recurrent Neural Network
by: Mamun, Nursadul, et al.
Published: (2026)
by: Mamun, Nursadul, et al.
Published: (2026)
Physics-Informed Neural Network for Volumetric Sound field Reconstruction of Speech Signals
by: Olivieri, Marco, et al.
Published: (2024)
by: Olivieri, Marco, et al.
Published: (2024)
Investigation of Deep Neural Network Acoustic Modelling Approaches for Low Resource Accented Mandarin Speech Recognition
by: Xie, Xurong, et al.
Published: (2022)
by: Xie, Xurong, et al.
Published: (2022)
Neural Speech Embeddings for Speech Synthesis Based on Deep Generative Networks
by: Lee, Seo-Hyun, et al.
Published: (2023)
by: Lee, Seo-Hyun, et al.
Published: (2023)
Leveraging Discriminative Latent Representations for Conditioning GAN-Based Speech Enhancement
by: Shetu, Shrishti Saha, et al.
Published: (2025)
by: Shetu, Shrishti Saha, et al.
Published: (2025)
Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis
by: Niu, Zhikang, et al.
Published: (2025)
by: Niu, Zhikang, et al.
Published: (2025)
Towards Frame-level Quality Predictions of Synthetic Speech
by: Kuhlmann, Michael, et al.
Published: (2025)
by: Kuhlmann, Michael, et al.
Published: (2025)
Universal Preference-Score-based Pairwise Speech Quality Assessment
by: Shi, Yu-Fei, et al.
Published: (2025)
by: Shi, Yu-Fei, et al.
Published: (2025)
Speech Synthesis From Continuous Features Using Per-Token Latent Diffusion
by: Turetzky, Arnon, et al.
Published: (2024)
by: Turetzky, Arnon, et al.
Published: (2024)
Latent-Domain Predictive Neural Speech Coding
by: Jiang, Xue, et al.
Published: (2022)
by: Jiang, Xue, et al.
Published: (2022)
Tackling Cognitive Impairment Detection from Speech: A submission to the PROCESS Challenge
by: Botelho, Catarina, et al.
Published: (2024)
by: Botelho, Catarina, et al.
Published: (2024)
HYFuse: Aligning Heterogeneous Speech Pre-Trained Representations in Hyperbolic Space for Speech Emotion Recognition
by: Phukan, Orchid Chetia, et al.
Published: (2025)
by: Phukan, Orchid Chetia, et al.
Published: (2025)
Simulating Native Speaker Shadowing for Nonnative Speech Assessment with Latent Speech Representations
by: Geng, Haopeng, et al.
Published: (2024)
by: Geng, Haopeng, et al.
Published: (2024)
CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis
by: Wu, Chun Yat, et al.
Published: (2025)
by: Wu, Chun Yat, et al.
Published: (2025)
Plugin Speech Enhancement: A Universal Speech Enhancement Framework Inspired by Dynamic Neural Network
by: Chen, Yanan, et al.
Published: (2024)
by: Chen, Yanan, et al.
Published: (2024)
Hybrid Real- And Complex-Valued Neural Network Concept For Low-Complexity Phase-Aware Speech Enhancement
by: Fiorio, Luan Vinícius, et al.
Published: (2025)
by: Fiorio, Luan Vinícius, et al.
Published: (2025)
Similar Items
-
Multivariate Probabilistic Assessment of Speech Quality
by: Cumlin, Fredrik, et al.
Published: (2025) -
Selection of Layers from Self-supervised Learning Models for Predicting Mean-Opinion-Score of Speech
by: Liang, Xinyu, et al.
Published: (2025) -
SA-SSL-MOS: Self-supervised Learning MOS Prediction with Spectral Augmentation for Generalized Multi-Rate Speech Assessment
by: Cao, Fengyuan, et al.
Published: (2026) -
Leveraging LLMs for Scalable Non-intrusive Speech Quality Assessment
by: Cumlin, Fredrik, et al.
Published: (2025) -
Rho-Perfect: Correlation Ceiling For Subjective Evaluation Datasets
by: Cumlin, Fredrik
Published: (2026)