Layer-wise Analysis for Quality of Multilingual Synthesized Speech
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cooper, Erica, Okamoto, Takuma, Ohtani, Yamato, Toda, Tomoki, Kawai, Hisashi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MOS-Bench: Benchmarking Generalization Abilities of Subjective Speech Quality Assessment Models
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
SHEET: A Multi-purpose Open-source Speech Human Evaluation Estimation Toolkit
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2025)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2025)
WaveNeXt 2: ConvNeXt-Based Fast Neural Vocoders With Residual Denoising and Sub-Modeling for GAN and Diffusion Models
von: Zhou, Wangzixi, et al.
Veröffentlicht: (2026)
von: Zhou, Wangzixi, et al.
Veröffentlicht: (2026)
The VoiceMOS Challenge 2024: Beyond Speech Quality Prediction
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
von: Hu, Cheng-Hung, et al.
Veröffentlicht: (2025)
von: Hu, Cheng-Hung, et al.
Veröffentlicht: (2025)
Investigation of perceptual music similarity focusing on each instrumental part
von: Hashizume, Yuka, et al.
Veröffentlicht: (2025)
von: Hashizume, Yuka, et al.
Veröffentlicht: (2025)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
Serial-OE: Anomalous sound detection based on serial method with outlier exposure capable of using small amounts of anomalous data for training
von: Kuroyanagi, Ibuki, et al.
Veröffentlicht: (2025)
von: Kuroyanagi, Ibuki, et al.
Veröffentlicht: (2025)
Eigenvoice Synthesis based on Model Editing for Speaker Generation
von: Murata, Masato, et al.
Veröffentlicht: (2025)
von: Murata, Masato, et al.
Veröffentlicht: (2025)
Improved Architecture for High-resolution Piano Transcription to Efficiently Capture Acoustic Characteristics of Music Signals
von: Mi, Jinyi, et al.
Veröffentlicht: (2024)
von: Mi, Jinyi, et al.
Veröffentlicht: (2024)
Improvements of Discriminative Feature Space Training for Anomalous Sound Detection in Unlabeled Conditions
von: Fujimura, Takuya, et al.
Veröffentlicht: (2024)
von: Fujimura, Takuya, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Speech Recognition Approach for Domain Challenges
von: Shen, Peng, et al.
Veröffentlicht: (2025)
von: Shen, Peng, et al.
Veröffentlicht: (2025)
QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on Autoregressive Moving Average Model
von: Chen, Shaowen, et al.
Veröffentlicht: (2025)
von: Chen, Shaowen, et al.
Veröffentlicht: (2025)
Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning
von: Ma, Ding, et al.
Veröffentlicht: (2026)
von: Ma, Ding, et al.
Veröffentlicht: (2026)
Electrolaryngeal Speech Intelligibility Enhancement Through Robust Linguistic Encoders
von: Violeta, Lester Phillip, et al.
Veröffentlicht: (2023)
von: Violeta, Lester Phillip, et al.
Veröffentlicht: (2023)
Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments
von: Yoneyama, Reo, et al.
Veröffentlicht: (2025)
von: Yoneyama, Reo, et al.
Veröffentlicht: (2025)
Diffusion Synthesizer for Efficient Multilingual Speech to Speech Translation
von: Hirschkind, Nameer, et al.
Veröffentlicht: (2024)
von: Hirschkind, Nameer, et al.
Veröffentlicht: (2024)
The AudioMOS Challenge 2025
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2025)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2025)
Serenade: A Singing Style Conversion Framework Based On Audio Infilling
von: Violeta, Lester Phillip, et al.
Veröffentlicht: (2025)
von: Violeta, Lester Phillip, et al.
Veröffentlicht: (2025)
Learning Separated Representations for Instrument-based Music Similarity
von: Hashizume, Yuka, et al.
Veröffentlicht: (2025)
von: Hashizume, Yuka, et al.
Veröffentlicht: (2025)
Improving Anomalous Sound Detection through Pseudo-anomalous Set Selection and Pseudo-label Utilization under Unlabeled Conditions
von: Kuroyanagi, Ibuki, et al.
Veröffentlicht: (2025)
von: Kuroyanagi, Ibuki, et al.
Veröffentlicht: (2025)
Learning Multidimensional Disentangled Representations of Instrumental Sounds for Musical Similarity Assessment
von: Hashizume, Yuka, et al.
Veröffentlicht: (2024)
von: Hashizume, Yuka, et al.
Veröffentlicht: (2024)
Wavehax: Aliasing-Free Neural Waveform Synthesis Based on 2D Convolution and Harmonic Prior for Reliable Complex Spectrogram Estimation
von: Yoneyama, Reo, et al.
Veröffentlicht: (2024)
von: Yoneyama, Reo, et al.
Veröffentlicht: (2024)
Prosody Labeling with Phoneme-BERT and Speech Foundation Models
von: Koriyama, Tomoki
Veröffentlicht: (2025)
von: Koriyama, Tomoki
Veröffentlicht: (2025)
CodecMOS-Accent: A MOS Benchmark of Resynthesized and TTS Speech from Neural Codecs Across English Accents
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2026)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2026)
A Comprehensive Study on the Effectiveness of ASR Representations for Noise-Robust Speech Emotion Recognition
von: Shi, Xiaohan, et al.
Veröffentlicht: (2023)
von: Shi, Xiaohan, et al.
Veröffentlicht: (2023)
Music Similarity Representation Learning Focusing on Individual Instruments with Source Separation and Human Preference
von: Imamura, Takehiro, et al.
Veröffentlicht: (2025)
von: Imamura, Takehiro, et al.
Veröffentlicht: (2025)
Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection
von: Kheir, Yassine El, et al.
Veröffentlicht: (2025)
von: Kheir, Yassine El, et al.
Veröffentlicht: (2025)
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
Interpreting End-to-End Deep Learning Models for Speech Source Localization Using Layer-wise Relevance Propagation
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
Two-stage Framework for Robust Speech Emotion Recognition Using Target Speaker Extraction in Human Speech Noise Conditions
von: Mi, Jinyi, et al.
Veröffentlicht: (2024)
von: Mi, Jinyi, et al.
Veröffentlicht: (2024)
An Extensive Analysis of the Singing Voice Conversion Challenge 2025 Evaluation Results
von: Violeta, Lester Phillip, et al.
Veröffentlicht: (2025)
von: Violeta, Lester Phillip, et al.
Veröffentlicht: (2025)
An Attribute Interpolation Method in Speech Synthesis by Model Merging
von: Murata, Masato, et al.
Veröffentlicht: (2024)
von: Murata, Masato, et al.
Veröffentlicht: (2024)
Efficient and Robust Long-Form Speech Recognition with Hybrid H3-Conformer
von: Honda, Tomoki, et al.
Veröffentlicht: (2024)
von: Honda, Tomoki, et al.
Veröffentlicht: (2024)
MF-AED-AEC: Speech Emotion Recognition by Leveraging Multimodal Fusion, Asr Error Detection, and Asr Error Correction
von: He, Jiajun, et al.
Veröffentlicht: (2024)
von: He, Jiajun, et al.
Veröffentlicht: (2024)
Frame-Wise Breath Detection with Self-Training: An Exploration of Enhancing Breath Naturalness in Text-to-Speech
von: Yang, Dong, et al.
Veröffentlicht: (2024)
von: Yang, Dong, et al.
Veröffentlicht: (2024)
Bayesian Speech Synthesizers Can Learn from Multiple Teachers
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
S2ST-Omni: Hierarchical Language-Aware SpeechLLM Adaptation for Multilingual Speech-to-Speech Translation
von: Pan, Yu, et al.
Veröffentlicht: (2025)
von: Pan, Yu, et al.
Veröffentlicht: (2025)
Handling Domain Shifts for Anomalous Sound Detection: A Review of DCASE-Related Work
von: Wilkinghoff, Kevin, et al.
Veröffentlicht: (2025)
von: Wilkinghoff, Kevin, et al.
Veröffentlicht: (2025)
Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MOS-Bench: Benchmarking Generalization Abilities of Subjective Speech Quality Assessment Models
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024) -
SHEET: A Multi-purpose Open-source Speech Human Evaluation Estimation Toolkit
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2025) -
WaveNeXt 2: ConvNeXt-Based Fast Neural Vocoders With Residual Denoising and Sub-Modeling for GAN and Diffusion Models
von: Zhou, Wangzixi, et al.
Veröffentlicht: (2026) -
The VoiceMOS Challenge 2024: Beyond Speech Quality Prediction
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024) -
Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
von: Hu, Cheng-Hung, et al.
Veröffentlicht: (2025)