SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Yu-Fei, Ai, Yang, Lu, Ye-Xin, Du, Hui-Peng, Ling, Zhen-Hua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
Pitch-and-Spectrum-Aware Singing Quality Assessment with Bias Correction and Model Fusion
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2024)
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2024)
Towards High-Quality and Efficient Speech Bandwidth Extension with Parallel Amplitude and Phase Prediction
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
APCodec: A Neural Audio Codec with Parallel Amplitude and Phase Spectrum Encoding and Decoding
von: Ai, Yang, et al.
Veröffentlicht: (2024)
von: Ai, Yang, et al.
Veröffentlicht: (2024)
Stage-Wise and Prior-Aware Neural Speech Phase Prediction
von: Liu, Fei, et al.
Veröffentlicht: (2024)
von: Liu, Fei, et al.
Veröffentlicht: (2024)
MDCTCodec: A Lightweight MDCT-based Neural Audio Codec towards High Sampling Rate and Low Bitrate Scenarios
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
Explicit Estimation of Magnitude and Phase Spectra in Parallel for High-Quality Speech Enhancement
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2023)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2023)
Low-Latency Neural Speech Phase Prediction based on Parallel Estimation Architecture and Anti-Wrapping Losses for Speech Generation Tasks
von: Ai, Yang, et al.
Veröffentlicht: (2024)
von: Ai, Yang, et al.
Veröffentlicht: (2024)
APCodec+: A Spectrum-Coding-Based High-Fidelity and High-Compression-Rate Neural Audio Codec with Staged Training Paradigm
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
Universal Preference-Score-based Pairwise Speech Quality Assessment
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2025)
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2025)
Vision-Integrated High-Quality Neural Speech Coding
von: Guo, Yao, et al.
Veröffentlicht: (2025)
von: Guo, Yao, et al.
Veröffentlicht: (2025)
Multi-Stage Speech Bandwidth Extension with Flexible Sampling Rate Control
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
SingMOS: An extensive Open-Source Singing Voice Dataset for MOS Prediction
von: Tang, Yuxun, et al.
Veröffentlicht: (2024)
von: Tang, Yuxun, et al.
Veröffentlicht: (2024)
HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement
von: Hussein, Amir, et al.
Veröffentlicht: (2025)
von: Hussein, Amir, et al.
Veröffentlicht: (2025)
The AudioMOS Challenge 2025
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2025)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2025)
CodecMOS-Accent: A MOS Benchmark of Resynthesized and TTS Speech from Neural Codecs Across English Accents
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2026)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2026)
Leveraging Prompt Learning and Pause Encoding for Alzheimer's Disease Detection
von: Liu, Yin-Long, et al.
Veröffentlicht: (2024)
von: Liu, Yin-Long, et al.
Veröffentlicht: (2024)
Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
von: Chung, Soo-Whan, et al.
Veröffentlicht: (2025)
von: Chung, Soo-Whan, et al.
Veröffentlicht: (2025)
Leveraging Self-supervised Audio Representations for Data-Efficient Acoustic Scene Classification
von: Cai, Yiqiang, et al.
Veröffentlicht: (2024)
von: Cai, Yiqiang, et al.
Veröffentlicht: (2024)
Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding
von: Zheng, Rui-Chen, et al.
Veröffentlicht: (2025)
von: Zheng, Rui-Chen, et al.
Veröffentlicht: (2025)
The VoiceMOS Challenge 2024: Beyond Speech Quality Prediction
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
A Neural Denoising Vocoder for Clean Waveform Generation from Noisy Mel-Spectrogram based on Amplitude and Phase Predictions
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
von: Chen, Wenxi, et al.
Veröffentlicht: (2025)
von: Chen, Wenxi, et al.
Veröffentlicht: (2025)
APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic Speech
von: Lian, Zhicheng, et al.
Veröffentlicht: (2025)
von: Lian, Zhicheng, et al.
Veröffentlicht: (2025)
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models
von: Zhang, Jing-Xuan, et al.
Veröffentlicht: (2025)
von: Zhang, Jing-Xuan, et al.
Veröffentlicht: (2025)
CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation
von: Deng, Ruifan, et al.
Veröffentlicht: (2025)
von: Deng, Ruifan, et al.
Veröffentlicht: (2025)
Sound Field Reconstruction Using a Compact Acoustics-informed Neural Network
von: Ma, Fei, et al.
Veröffentlicht: (2024)
von: Ma, Fei, et al.
Veröffentlicht: (2024)
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
von: Du, Chenpeng, et al.
Veröffentlicht: (2022)
von: Du, Chenpeng, et al.
Veröffentlicht: (2022)
A Distilled Low-Latency Neural Vocoder with Explicit Amplitude and Phase Prediction
von: Du, Hui-Peng, et al.
Veröffentlicht: (2025)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2025)
Investigating the Reasonable Effectiveness of Speaker Pre-Trained Models and their Synergistic Power for SingMOS Prediction
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction
von: Emon, Jakaria Islam, et al.
Veröffentlicht: (2025)
von: Emon, Jakaria Islam, et al.
Veröffentlicht: (2025)
Selecting N-lowest scores for training MOS prediction models
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
Multi-Sampling-Frequency Naturalness MOS Prediction Using Self-Supervised Learning Model with Sampling-Frequency-Independent Layer
von: Nishikawa, Go, et al.
Veröffentlicht: (2025)
von: Nishikawa, Go, et al.
Veröffentlicht: (2025)
The T12 System for AudioMOS Challenge 2025: Audio Aesthetics Score Prediction System Using KAN- and VERSA-based Models
von: Yamamoto, Katsuhiko, et al.
Veröffentlicht: (2025)
von: Yamamoto, Katsuhiko, et al.
Veröffentlicht: (2025)
ASTAR-NTU solution to AudioMOS Challenge 2025 Track1
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025)
Abusive Speech Detection in Indic Languages Using Acoustic Features
von: Spiesberger, Anika A., et al.
Veröffentlicht: (2024)
von: Spiesberger, Anika A., et al.
Veröffentlicht: (2024)
XANE: eXplainable Acoustic Neural Embeddings
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024) -
Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025) -
Pitch-and-Spectrum-Aware Singing Quality Assessment with Bias Correction and Model Fusion
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2024) -
Towards High-Quality and Efficient Speech Bandwidth Extension with Parallel Amplitude and Phase Prediction
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024) -
ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)