Selecting N-lowest scores for training MOS prediction models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kondo, Yuto, Kameoka, Hirokazu, Tanaka, Kou, Kaneko, Takuhiro |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LatentVoiceGrad: Nonparallel Voice Conversion with Latent Diffusion/Flow-Matching Models
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2025)
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2025)
Rethinking Mean Opinion Scores in Speech Quality Assessment: Aggregation through Quantized Distribution Fitting
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
Vocoder-Projected Feature Discriminator
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025)
FasterVoiceGrad: Faster One-step Diffusion-Based Voice Conversion with Adversarial Diffusion Conversion Distillation
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025)
MeanVoiceFlow: One-step Nonparallel Voice Conversion with Mean Flows
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2026)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2026)
FastVoiceGrad: One-step Diffusion-Based Voice Conversion with Adversarial Conditional Diffusion Distillation
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024)
Training Generative Adversarial Network-Based Vocoder with Limited Data Using Augmentation-Conditional Discriminator
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024)
VoiceGrad: Non-Parallel Any-to-Many Voice Conversion with Annealed Langevin Dynamics
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2020)
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2020)
GE2E-AC: Generalized End-to-End Loss Training for Accent Classification
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2024)
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2024)
SingMOS: An extensive Open-Source Singing Voice Dataset for MOS Prediction
von: Tang, Yuxun, et al.
Veröffentlicht: (2024)
von: Tang, Yuxun, et al.
Veröffentlicht: (2024)
The AudioMOS Challenge 2025
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2025)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2025)
CodecMOS-Accent: A MOS Benchmark of Resynthesized and TTS Speech from Neural Codecs Across English Accents
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2026)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2026)
The VoiceMOS Challenge 2024: Beyond Speech Quality Prediction
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
ASTAR-NTU solution to AudioMOS Challenge 2025 Track1
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025)
MOS-Bench: Benchmarking Generalization Abilities of Subjective Speech Quality Assessment Models
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2024)
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2024)
Analysing the Masked predictive coding training criterion for pre-training a Speech Representation Model
von: Yadav, Hemant, et al.
Veröffentlicht: (2023)
von: Yadav, Hemant, et al.
Veröffentlicht: (2023)
Investigating the Reasonable Effectiveness of Speaker Pre-Trained Models and their Synergistic Power for SingMOS Prediction
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic Speech
von: Lian, Zhicheng, et al.
Veröffentlicht: (2025)
von: Lian, Zhicheng, et al.
Veröffentlicht: (2025)
Quality Assessment of Noisy and Enhanced Speech with Limited Data: UWB-NTIS System for VoiceMOS 2024
von: Kunešová, Marie, et al.
Veröffentlicht: (2025)
von: Kunešová, Marie, et al.
Veröffentlicht: (2025)
Learning to assess subjective impressions from speech
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
Multi-Sampling-Frequency Naturalness MOS Prediction Using Self-Supervised Learning Model with Sampling-Frequency-Independent Layer
von: Nishikawa, Go, et al.
Veröffentlicht: (2025)
von: Nishikawa, Go, et al.
Veröffentlicht: (2025)
The T12 System for AudioMOS Challenge 2025: Audio Aesthetics Score Prediction System Using KAN- and VERSA-based Models
von: Yamamoto, Katsuhiko, et al.
Veröffentlicht: (2025)
von: Yamamoto, Katsuhiko, et al.
Veröffentlicht: (2025)
A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models
von: Whetten, Ryan, et al.
Veröffentlicht: (2026)
von: Whetten, Ryan, et al.
Veröffentlicht: (2026)
Bridging the gap between training and inference in LM-based TTS models
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
Reconstructing the Charlie Parker Omnibook using an audio-to-score automatic transcription pipeline
von: Riley, Xavier, et al.
Veröffentlicht: (2024)
von: Riley, Xavier, et al.
Veröffentlicht: (2024)
Using RLHF to align speech enhancement approaches to mean-opinion quality scores
von: Kumar, Anurag, et al.
Veröffentlicht: (2024)
von: Kumar, Anurag, et al.
Veröffentlicht: (2024)
SingMOS-Pro: An Comprehensive Benchmark for Singing Quality Assessment
von: Tang, Yuxun, et al.
Veröffentlicht: (2025)
von: Tang, Yuxun, et al.
Veröffentlicht: (2025)
Real-time Speech Extraction Using Spatially Regularized Independent Low-rank Matrix Analysis and Rank-constrained Spatial Covariance Matrix Estimation
von: Ishikawa, Yuto, et al.
Veröffentlicht: (2024)
von: Ishikawa, Yuto, et al.
Veröffentlicht: (2024)
HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis
von: Nishimura, Yuto, et al.
Veröffentlicht: (2024)
von: Nishimura, Yuto, et al.
Veröffentlicht: (2024)
Misophonia Trigger Sound Detection on Synthetic Soundscapes Using a Hybrid Model with a Frozen Pre-Trained CNN and a Time-Series Module
von: Sashida, Kurumi, et al.
Veröffentlicht: (2026)
von: Sashida, Kurumi, et al.
Veröffentlicht: (2026)
An interpretable speech foundation model for depression detection by revealing prediction-relevant acoustic features from long speech
von: Deng, Qingkun, et al.
Veröffentlicht: (2024)
von: Deng, Qingkun, et al.
Veröffentlicht: (2024)
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction
von: Emon, Jakaria Islam, et al.
Veröffentlicht: (2025)
von: Emon, Jakaria Islam, et al.
Veröffentlicht: (2025)
The TMU System for the XACLE Challenge: Training Large Audio Language Models with CLAP Pseudo-Labels
von: Tsutsumi, Ayuto, et al.
Veröffentlicht: (2026)
von: Tsutsumi, Ayuto, et al.
Veröffentlicht: (2026)
Consistency Based Unsupervised Self-training For ASR Personalisation
von: Zhang, Jisi, et al.
Veröffentlicht: (2024)
von: Zhang, Jisi, et al.
Veröffentlicht: (2024)
Event Classification by Physics-informed Inpainting for Distributed Multichannel Acoustic Sensor with Partially Degraded Channels
von: Tonami, Noriyuki, et al.
Veröffentlicht: (2026)
von: Tonami, Noriyuki, et al.
Veröffentlicht: (2026)
Trainingless Adaptation of Pretrained Models for Environmental Sound Classification
von: Tonami, Noriyuki, et al.
Veröffentlicht: (2024)
von: Tonami, Noriyuki, et al.
Veröffentlicht: (2024)
LLM supervised Pre-training for Multimodal Emotion Recognition in Conversations
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
Speaker Disentanglement of Speech Pre-trained Model Based on Interpretability
von: Zhu, Xiaoxu, et al.
Veröffentlicht: (2025)
von: Zhu, Xiaoxu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LatentVoiceGrad: Nonparallel Voice Conversion with Latent Diffusion/Flow-Matching Models
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2025) -
Rethinking Mean Opinion Scores in Speech Quality Assessment: Aggregation through Quantized Distribution Fitting
von: Kondo, Yuto, et al.
Veröffentlicht: (2025) -
JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles
von: Kondo, Yuto, et al.
Veröffentlicht: (2025) -
Vocoder-Projected Feature Discriminator
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025) -
FasterVoiceGrad: Faster One-step Diffusion-Based Voice Conversion with Adversarial Diffusion Conversion Distillation
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025)