Selection of Layers from Self-supervised Learning Models for Predicting Mean-Opinion-Score of Speech
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liang, Xinyu, Cumlin, Fredrik, Ungureanu, Victor, Reddy, Chandan K. A., Schuldt, Christian, Chatterjee, Saikat |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SA-SSL-MOS: Self-supervised Learning MOS Prediction with Spectral Augmentation for Generalized Multi-Rate Speech Assessment
von: Cao, Fengyuan, et al.
Veröffentlicht: (2026)
von: Cao, Fengyuan, et al.
Veröffentlicht: (2026)
Multivariate Probabilistic Assessment of Speech Quality
von: Cumlin, Fredrik, et al.
Veröffentlicht: (2025)
von: Cumlin, Fredrik, et al.
Veröffentlicht: (2025)
Impairments are Clustered in Latents of Deep Neural Network-based Speech Quality Models
von: Cumlin, Fredrik, et al.
Veröffentlicht: (2025)
von: Cumlin, Fredrik, et al.
Veröffentlicht: (2025)
Leveraging LLMs for Scalable Non-intrusive Speech Quality Assessment
von: Cumlin, Fredrik, et al.
Veröffentlicht: (2025)
von: Cumlin, Fredrik, et al.
Veröffentlicht: (2025)
Rethinking Mean Opinion Scores in Speech Quality Assessment: Aggregation through Quantized Distribution Fitting
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
Uncertainty-Aware Mean Opinion Score Prediction
von: Wang, Hui, et al.
Veröffentlicht: (2024)
von: Wang, Hui, et al.
Veröffentlicht: (2024)
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024)
MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
Probing Self-supervised Learning Models with Target Speech Extraction
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
Target Speech Extraction with Pre-trained Self-supervised Learning Models
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
Multi-Channel MOSRA: Mean Opinion Score and Room Acoustics Estimation Using Simulated Data and a Teacher Model
von: Coldenhoff, Jozef, et al.
Veröffentlicht: (2023)
von: Coldenhoff, Jozef, et al.
Veröffentlicht: (2023)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing Tasks
von: Huang, He, et al.
Veröffentlicht: (2024)
von: Huang, He, et al.
Veröffentlicht: (2024)
Causal Self-supervised Pretrained Frontend with Predictive Code for Speech Separation
von: Wang, Wupeng, et al.
Veröffentlicht: (2025)
von: Wang, Wupeng, et al.
Veröffentlicht: (2025)
Less Forgetting for Better Generalization: Exploring Continual-learning Fine-tuning Methods for Speech Self-supervised Representations
von: Zaiem, Salah, et al.
Veröffentlicht: (2024)
von: Zaiem, Salah, et al.
Veröffentlicht: (2024)
STONE: Self-supervised Tonality Estimator
von: Kong, Yuexuan, et al.
Veröffentlicht: (2024)
von: Kong, Yuexuan, et al.
Veröffentlicht: (2024)
MeanFlowSE: One-Step Generative Speech Enhancement via MeanFlow
von: Zhu, Yike, et al.
Veröffentlicht: (2025)
von: Zhu, Yike, et al.
Veröffentlicht: (2025)
Combined Generative and Predictive Modeling for Speech Super-resolution
von: Wang, Heming, et al.
Veröffentlicht: (2024)
von: Wang, Heming, et al.
Veröffentlicht: (2024)
Rho-Perfect: Correlation Ceiling For Subjective Evaluation Datasets
von: Cumlin, Fredrik
Veröffentlicht: (2026)
von: Cumlin, Fredrik
Veröffentlicht: (2026)
A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models
von: Whetten, Ryan, et al.
Veröffentlicht: (2026)
von: Whetten, Ryan, et al.
Veröffentlicht: (2026)
SuPseudo: A Pseudo-supervised Learning Method for Neural Speech Enhancement in Far-field Speech Recognition
von: Luo, Longjie, et al.
Veröffentlicht: (2025)
von: Luo, Longjie, et al.
Veröffentlicht: (2025)
Self-supervised Multimodal Speech Representations for the Assessment of Schizophrenia Symptoms
von: Premananth, Gowtham, et al.
Veröffentlicht: (2024)
von: Premananth, Gowtham, et al.
Veröffentlicht: (2024)
Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations
von: Meghanani, Amit, et al.
Veröffentlicht: (2026)
von: Meghanani, Amit, et al.
Veröffentlicht: (2026)
Universal Score-based Speech Enhancement with High Content Preservation
von: Scheibler, Robin, et al.
Veröffentlicht: (2024)
von: Scheibler, Robin, et al.
Veröffentlicht: (2024)
Universal Preference-Score-based Pairwise Speech Quality Assessment
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2025)
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2025)
Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning
von: Zhu, Xinfa, et al.
Veröffentlicht: (2023)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2023)
Towards Sub-millisecond Latency Real-Time Speech Enhancement Models on Hearables
von: Dementyev, Artem, et al.
Veröffentlicht: (2024)
von: Dementyev, Artem, et al.
Veröffentlicht: (2024)
Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
von: Hu, Cheng-Hung, et al.
Veröffentlicht: (2025)
von: Hu, Cheng-Hung, et al.
Veröffentlicht: (2025)
Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)
Layer-wise Analysis for Quality of Multilingual Synthesized Speech
von: Cooper, Erica, et al.
Veröffentlicht: (2025)
von: Cooper, Erica, et al.
Veröffentlicht: (2025)
Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology
von: Moell, Birger, et al.
Veröffentlicht: (2025)
von: Moell, Birger, et al.
Veröffentlicht: (2025)
PESTO: Pitch Estimation with Self-supervised Transposition-equivariant Objective
von: Riou, Alain, et al.
Veröffentlicht: (2023)
von: Riou, Alain, et al.
Veröffentlicht: (2023)
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
Self-supervised Reflective Learning through Self-distillation and Online Clustering for Speaker Representation Learning
von: Cai, Danwei, et al.
Veröffentlicht: (2024)
von: Cai, Danwei, et al.
Veröffentlicht: (2024)
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
von: Li, Jialu, et al.
Veröffentlicht: (2024)
von: Li, Jialu, et al.
Veröffentlicht: (2024)
No Audiogram: Leveraging Existing Scores for Personalized Speech Intelligibility Prediction
von: Zhou, Haoshuai, et al.
Veröffentlicht: (2025)
von: Zhou, Haoshuai, et al.
Veröffentlicht: (2025)
Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)
Self-supervised Speech Representations Still Struggle with African American Vernacular English
von: Chang, Kalvin, et al.
Veröffentlicht: (2024)
von: Chang, Kalvin, et al.
Veröffentlicht: (2024)
SSHR: Leveraging Self-supervised Hierarchical Representations for Multilingual Automatic Speech Recognition
von: Xue, Hongfei, et al.
Veröffentlicht: (2023)
von: Xue, Hongfei, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
SA-SSL-MOS: Self-supervised Learning MOS Prediction with Spectral Augmentation for Generalized Multi-Rate Speech Assessment
von: Cao, Fengyuan, et al.
Veröffentlicht: (2026) -
Multivariate Probabilistic Assessment of Speech Quality
von: Cumlin, Fredrik, et al.
Veröffentlicht: (2025) -
Impairments are Clustered in Latents of Deep Neural Network-based Speech Quality Models
von: Cumlin, Fredrik, et al.
Veröffentlicht: (2025) -
Leveraging LLMs for Scalable Non-intrusive Speech Quality Assessment
von: Cumlin, Fredrik, et al.
Veröffentlicht: (2025) -
Rethinking Mean Opinion Scores in Speech Quality Assessment: Aggregation through Quantized Distribution Fitting
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)