Good practices for evaluation of synthesized speech
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cooper, Erica, Maguer, Sébastien Le, Klabbers, Esther, Yamagishi, Junichi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards An Integrated Approach for Expressive Piano Performance Synthesis from Music Scores
von: Tang, Jingjing, et al.
Veröffentlicht: (2025)
von: Tang, Jingjing, et al.
Veröffentlicht: (2025)
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
Spoofing-Aware Speaker Verification Robust Against Domain and Channel Mismatches
von: Zeng, Chang, et al.
Veröffentlicht: (2024)
von: Zeng, Chang, et al.
Veröffentlicht: (2024)
Towards Data Drift Monitoring for Speech Deepfake Detection in the context of MLOps
von: Wang, Xin, et al.
Veröffentlicht: (2025)
von: Wang, Xin, et al.
Veröffentlicht: (2025)
FakeMark: Deepfake Speech Attribution With Watermarked Artifacts
von: Ge, Wanying, et al.
Veröffentlicht: (2025)
von: Ge, Wanying, et al.
Veröffentlicht: (2025)
Does Fine-tuning by Reinforcement Learning Improve Generalization in Binary Speech Deepfake Detection?
von: Wang, Xin, et al.
Veröffentlicht: (2026)
von: Wang, Xin, et al.
Veröffentlicht: (2026)
VoxEffects: A Speech-Oriented Audio Effects Dataset and Benchmark
von: Zhang, Zhe, et al.
Veröffentlicht: (2026)
von: Zhang, Zhe, et al.
Veröffentlicht: (2026)
AfriHuBERT: A self-supervised speech representation model for African languages
von: Alabi, Jesujoba O., et al.
Veröffentlicht: (2024)
von: Alabi, Jesujoba O., et al.
Veröffentlicht: (2024)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
Improving curriculum learning for target speaker extraction with synthetic speakers
von: Liu, Yun, et al.
Veröffentlicht: (2024)
von: Liu, Yun, et al.
Veröffentlicht: (2024)
Post-training for Deepfake Speech Detection
von: Ge, Wanying, et al.
Veröffentlicht: (2025)
von: Ge, Wanying, et al.
Veröffentlicht: (2025)
Spoof Diarization: "What Spoofed When" in Partially Spoofed Audio
von: Zhang, Lin, et al.
Veröffentlicht: (2024)
von: Zhang, Lin, et al.
Veröffentlicht: (2024)
The VoiceMOS Challenge 2024: Beyond Speech Quality Prediction
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
Target Speaker Extraction with Curriculum Learning
von: Liu, Yun, et al.
Veröffentlicht: (2024)
von: Liu, Yun, et al.
Veröffentlicht: (2024)
Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data
von: Liu, Yun, et al.
Veröffentlicht: (2024)
von: Liu, Yun, et al.
Veröffentlicht: (2024)
Explaining Speaker and Spoof Embeddings via Probing
von: Liu, Xuechen, et al.
Veröffentlicht: (2024)
von: Liu, Xuechen, et al.
Veröffentlicht: (2024)
Human perception of audio deepfakes: the role of language and speaking style
von: Segundo, Eugenia San, et al.
Veröffentlicht: (2025)
von: Segundo, Eugenia San, et al.
Veröffentlicht: (2025)
Assessing speech quality metrics for evaluation of neural audio codecs under clean speech conditions
von: Mack, Wolfgang, et al.
Veröffentlicht: (2025)
von: Mack, Wolfgang, et al.
Veröffentlicht: (2025)
Quantifying Source Speaker Leakage in One-to-One Voice Conversion
von: Wellington, Scott, et al.
Veröffentlicht: (2025)
von: Wellington, Scott, et al.
Veröffentlicht: (2025)
A Preliminary Case Study on Long-Form In-the-Wild Audio Spoofing Detection
von: Liu, Xuechen, et al.
Veröffentlicht: (2024)
von: Liu, Xuechen, et al.
Veröffentlicht: (2024)
From Sharpness to Better Generalization for Speech Deepfake Detection
von: Huang, Wen, et al.
Veröffentlicht: (2025)
von: Huang, Wen, et al.
Veröffentlicht: (2025)
Mitigating Language Mismatch in SSL-Based Speaker Anonymization
von: Zhang, Zhe, et al.
Veröffentlicht: (2025)
von: Zhang, Zhe, et al.
Veröffentlicht: (2025)
Revisiting and Improving Scoring Fusion for Spoofing-aware Speaker Verification Using Compositional Data Analysis
von: Wang, Xin, et al.
Veröffentlicht: (2024)
von: Wang, Xin, et al.
Veröffentlicht: (2024)
Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
LENS-DF: Deepfake Detection and Temporal Localization for Long-Form Noisy Speech
von: Liu, Xuechen, et al.
Veröffentlicht: (2025)
von: Liu, Xuechen, et al.
Veröffentlicht: (2025)
An Initial Investigation of Language Adaptation for TTS Systems under Low-resource Scenarios
von: Gong, Cheng, et al.
Veröffentlicht: (2024)
von: Gong, Cheng, et al.
Veröffentlicht: (2024)
Deepfake Word Detection by Next-token Prediction using Fine-tuned Whisper
von: Tran, Hoan My, et al.
Veröffentlicht: (2026)
von: Tran, Hoan My, et al.
Veröffentlicht: (2026)
MUSHRA-1S: A scalable and sensitive test approach for evaluating top-tier speech processing systems
von: Lechler, Laura, et al.
Veröffentlicht: (2025)
von: Lechler, Laura, et al.
Veröffentlicht: (2025)
ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech
von: Wang, Xin, et al.
Veröffentlicht: (2025)
von: Wang, Xin, et al.
Veröffentlicht: (2025)
The VoicePrivacy 2022 Challenge: Progress and Perspectives in Voice Anonymisation
von: Panariello, Michele, et al.
Veröffentlicht: (2024)
von: Panariello, Michele, et al.
Veröffentlicht: (2024)
The First VoicePrivacy Attacker Challenge
von: Tomashenko, Natalia, et al.
Veröffentlicht: (2025)
von: Tomashenko, Natalia, et al.
Veröffentlicht: (2025)
The First VoicePrivacy Attacker Challenge Evaluation Plan
von: Tomashenko, Natalia, et al.
Veröffentlicht: (2024)
von: Tomashenko, Natalia, et al.
Veröffentlicht: (2024)
Lightweight speech enhancement guided target speech extraction in noisy multi-speaker scenarios
von: Huang, Ziling, et al.
Veröffentlicht: (2025)
von: Huang, Ziling, et al.
Veröffentlicht: (2025)
Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
von: Saon, George, et al.
Veröffentlicht: (2025)
von: Saon, George, et al.
Veröffentlicht: (2025)
SLM-S2ST: A multimodal language model for direct speech-to-speech translation
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
Ensemble of classifiers for speech evaluation
von: Belokrylov, G., et al.
Veröffentlicht: (2024)
von: Belokrylov, G., et al.
Veröffentlicht: (2024)
Spoofing attack augmentation: can differently-trained attack models improve generalisation?
von: Ge, Wanying, et al.
Veröffentlicht: (2023)
von: Ge, Wanying, et al.
Veröffentlicht: (2023)
SHEET: A Multi-purpose Open-source Speech Human Evaluation Estimation Toolkit
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2025)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2025)
MOS-Bench: Benchmarking Generalization Abilities of Subjective Speech Quality Assessment Models
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
CodecMOS-Accent: A MOS Benchmark of Resynthesized and TTS Speech from Neural Codecs Across English Accents
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2026)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Towards An Integrated Approach for Expressive Piano Performance Synthesis from Music Scores
von: Tang, Jingjing, et al.
Veröffentlicht: (2025) -
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024) -
Spoofing-Aware Speaker Verification Robust Against Domain and Channel Mismatches
von: Zeng, Chang, et al.
Veröffentlicht: (2024) -
Towards Data Drift Monitoring for Speech Deepfake Detection in the context of MLOps
von: Wang, Xin, et al.
Veröffentlicht: (2025) -
FakeMark: Deepfake Speech Attribution With Watermarked Artifacts
von: Ge, Wanying, et al.
Veröffentlicht: (2025)