Universal Preference-Score-based Pairwise Speech Quality Assessment
Fuente:
arXiv
Guardado en:
| Autores principales: | Shi, Yu-Fei, Ai, Yang, Ling, Zhen-Hua |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Low-Latency Neural Speech Phase Prediction based on Parallel Estimation Architecture and Anti-Wrapping Losses for Speech Generation Tasks
por: Ai, Yang, et al.
Publicado: (2024)
por: Ai, Yang, et al.
Publicado: (2024)
Pitch-and-Spectrum-Aware Singing Quality Assessment with Bias Correction and Model Fusion
por: Shi, Yu-Fei, et al.
Publicado: (2024)
por: Shi, Yu-Fei, et al.
Publicado: (2024)
Explicit Estimation of Magnitude and Phase Spectra in Parallel for High-Quality Speech Enhancement
por: Lu, Ye-Xin, et al.
Publicado: (2023)
por: Lu, Ye-Xin, et al.
Publicado: (2023)
Vision-Integrated High-Quality Neural Speech Coding
por: Guo, Yao, et al.
Publicado: (2025)
por: Guo, Yao, et al.
Publicado: (2025)
Towards High-Quality and Efficient Speech Bandwidth Extension with Parallel Amplitude and Phase Prediction
por: Lu, Ye-Xin, et al.
Publicado: (2024)
por: Lu, Ye-Xin, et al.
Publicado: (2024)
Multi-Stage Speech Bandwidth Extension with Flexible Sampling Rate Control
por: Lu, Ye-Xin, et al.
Publicado: (2024)
por: Lu, Ye-Xin, et al.
Publicado: (2024)
SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features
por: Shi, Yu-Fei, et al.
Publicado: (2024)
por: Shi, Yu-Fei, et al.
Publicado: (2024)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
por: Lu, Ye-Xin, et al.
Publicado: (2025)
por: Lu, Ye-Xin, et al.
Publicado: (2025)
Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising
por: Lu, Ye-Xin, et al.
Publicado: (2025)
por: Lu, Ye-Xin, et al.
Publicado: (2025)
Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
por: Hu, Cheng-Hung, et al.
Publicado: (2025)
por: Hu, Cheng-Hung, et al.
Publicado: (2025)
Universal Score-based Speech Enhancement with High Content Preservation
por: Scheibler, Robin, et al.
Publicado: (2024)
por: Scheibler, Robin, et al.
Publicado: (2024)
Rethinking Mean Opinion Scores in Speech Quality Assessment: Aggregation through Quantized Distribution Fitting
por: Kondo, Yuto, et al.
Publicado: (2025)
por: Kondo, Yuto, et al.
Publicado: (2025)
BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation
por: Du, Hui-Peng, et al.
Publicado: (2024)
por: Du, Hui-Peng, et al.
Publicado: (2024)
APCodec+: A Spectrum-Coding-Based High-Fidelity and High-Compression-Rate Neural Audio Codec with Staged Training Paradigm
por: Du, Hui-Peng, et al.
Publicado: (2024)
por: Du, Hui-Peng, et al.
Publicado: (2024)
Stage-Wise and Prior-Aware Neural Speech Phase Prediction
por: Liu, Fei, et al.
Publicado: (2024)
por: Liu, Fei, et al.
Publicado: (2024)
SCOREQ: Speech Quality Assessment with Contrastive Regression
por: Ragano, Alessandro, et al.
Publicado: (2024)
por: Ragano, Alessandro, et al.
Publicado: (2024)
Improving Speech Enhancement with Multi-Metric Supervision from Learned Quality Assessment
por: Wang, Wei, et al.
Publicado: (2025)
por: Wang, Wei, et al.
Publicado: (2025)
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models
por: Zhang, Jing-Xuan, et al.
Publicado: (2025)
por: Zhang, Jing-Xuan, et al.
Publicado: (2025)
Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding
por: Zheng, Rui-Chen, et al.
Publicado: (2025)
por: Zheng, Rui-Chen, et al.
Publicado: (2025)
Self-Supervised Speech Quality Assessment (S3QA): Leveraging Speech Foundation Models for a Scalable Speech Quality Metric
por: Ogg, Mattson, et al.
Publicado: (2025)
por: Ogg, Mattson, et al.
Publicado: (2025)
MDCTCodec: A Lightweight MDCT-based Neural Audio Codec towards High Sampling Rate and Low Bitrate Scenarios
por: Jiang, Xiao-Hang, et al.
Publicado: (2024)
por: Jiang, Xiao-Hang, et al.
Publicado: (2024)
ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram
por: Jiang, Xiao-Hang, et al.
Publicado: (2024)
por: Jiang, Xiao-Hang, et al.
Publicado: (2024)
APCodec: A Neural Audio Codec with Parallel Amplitude and Phase Spectrum Encoding and Decoding
por: Ai, Yang, et al.
Publicado: (2024)
por: Ai, Yang, et al.
Publicado: (2024)
ReFESS-QI: Reference-Free Evaluation For Speech Separation With Joint Quality And Intelligibility Scoring
por: Frummer, Ari, et al.
Publicado: (2025)
por: Frummer, Ari, et al.
Publicado: (2025)
MambaRate: Speech Quality Assessment Across Different Sampling Rates
por: Kakoulidis, Panos, et al.
Publicado: (2025)
por: Kakoulidis, Panos, et al.
Publicado: (2025)
Refining Self-Supervised Learnt Speech Representation using Brain Activations
por: Li, Hengyu, et al.
Publicado: (2024)
por: Li, Hengyu, et al.
Publicado: (2024)
Fine-grained Preference Optimization Improves Zero-shot Text-to-Speech
por: Yao, Jixun, et al.
Publicado: (2025)
por: Yao, Jixun, et al.
Publicado: (2025)
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
por: Xia, Kangxiang, et al.
Publicado: (2025)
por: Xia, Kangxiang, et al.
Publicado: (2025)
Distillation and Pruning for Scalable Self-Supervised Representation-Based Speech Quality Assessment
por: Stahl, Benjamin, et al.
Publicado: (2025)
por: Stahl, Benjamin, et al.
Publicado: (2025)
Advancing Speech Quality Assessment Through Scientific Challenges and Open-source Activities
por: Huang, Wen-Chin
Publicado: (2025)
por: Huang, Wen-Chin
Publicado: (2025)
MOS-Bench: Benchmarking Generalization Abilities of Subjective Speech Quality Assessment Models
por: Huang, Wen-Chin, et al.
Publicado: (2024)
por: Huang, Wen-Chin, et al.
Publicado: (2024)
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
por: Li, Xuyuan, et al.
Publicado: (2024)
por: Li, Xuyuan, et al.
Publicado: (2024)
SECodec: Structural Entropy-based Compressive Speech Representation Codec for Speech Language Models
por: Wang, Linqin, et al.
Publicado: (2024)
por: Wang, Linqin, et al.
Publicado: (2024)
Uni-VERSA: Versatile Speech Assessment with a Unified Network
por: Shi, Jiatong, et al.
Publicado: (2025)
por: Shi, Jiatong, et al.
Publicado: (2025)
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
por: Wang, Hui, et al.
Publicado: (2025)
por: Wang, Hui, et al.
Publicado: (2025)
A Study on Incorporating Whisper for Robust Speech Assessment
por: Zezario, Ryandhimas E., et al.
Publicado: (2023)
por: Zezario, Ryandhimas E., et al.
Publicado: (2023)
Universal Speech Enhancement with Regression and Generative Mamba
por: Chao, Rong, et al.
Publicado: (2025)
por: Chao, Rong, et al.
Publicado: (2025)
The Universal Personalizer: Few-Shot Dysarthric Speech Recognition via Meta-Learning
por: Agarwal, Dhruuv, et al.
Publicado: (2025)
por: Agarwal, Dhruuv, et al.
Publicado: (2025)
Quality Assessment of Noisy and Enhanced Speech with Limited Data: UWB-NTIS System for VoiceMOS 2024
por: Kunešová, Marie, et al.
Publicado: (2025)
por: Kunešová, Marie, et al.
Publicado: (2025)
NOMAD: Unsupervised Learning of Perceptual Embeddings for Speech Enhancement and Non-matching Reference Audio Quality Assessment
por: Ragano, Alessandro, et al.
Publicado: (2023)
por: Ragano, Alessandro, et al.
Publicado: (2023)
Ejemplares similares
-
Low-Latency Neural Speech Phase Prediction based on Parallel Estimation Architecture and Anti-Wrapping Losses for Speech Generation Tasks
por: Ai, Yang, et al.
Publicado: (2024) -
Pitch-and-Spectrum-Aware Singing Quality Assessment with Bias Correction and Model Fusion
por: Shi, Yu-Fei, et al.
Publicado: (2024) -
Explicit Estimation of Magnitude and Phase Spectra in Parallel for High-Quality Speech Enhancement
por: Lu, Ye-Xin, et al.
Publicado: (2023) -
Vision-Integrated High-Quality Neural Speech Coding
por: Guo, Yao, et al.
Publicado: (2025) -
Towards High-Quality and Efficient Speech Bandwidth Extension with Parallel Amplitude and Phase Prediction
por: Lu, Ye-Xin, et al.
Publicado: (2024)