The ISCSLP 2024 Conversational Voice Clone (CoVoC) Challenge: Tasks, Results and Findings
Fuente:
arXiv
Salvato in:
| Autori principali: | Xia, Kangxiang, Guo, Dake, Yao, Jixun, Xue, Liumeng, Li, Hanzhao, Wang, Shuai, Guo, Zhao, Xie, Lei, Zhang, Qingqing, Luo, Lei, Dong, Minghui, Sun, Peng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The NPU-HWC System for the ISCSLP 2024 Inspirational and Convincing Audio Generation Challenge
di: Guo, Dake, et al.
Pubblicazione: (2024)
di: Guo, Dake, et al.
Pubblicazione: (2024)
The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024
di: Zhou, Shuoyi, et al.
Pubblicazione: (2024)
di: Zhou, Shuoyi, et al.
Pubblicazione: (2024)
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
di: Xia, Kangxiang, et al.
Pubblicazione: (2025)
di: Xia, Kangxiang, et al.
Pubblicazione: (2025)
StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding
di: Guo, Dake, et al.
Pubblicazione: (2025)
di: Guo, Dake, et al.
Pubblicazione: (2025)
SponTTS: modeling and transferring spontaneous style for TTS
di: Li, Hanzhao, et al.
Pubblicazione: (2023)
di: Li, Hanzhao, et al.
Pubblicazione: (2023)
KALL-E:Autoregressive Speech Synthesis with Next-Distribution Prediction
di: Xia, Kangxiang, et al.
Pubblicazione: (2024)
di: Xia, Kangxiang, et al.
Pubblicazione: (2024)
Multi-level Temporal-channel Speaker Retrieval for Zero-shot Voice Conversion
di: Wang, Zhichao, et al.
Pubblicazione: (2023)
di: Wang, Zhichao, et al.
Pubblicazione: (2023)
DualVC 2: Dynamic Masked Convolution for Unified Streaming and Non-Streaming Voice Conversion
di: Ning, Ziqian, et al.
Pubblicazione: (2023)
di: Ning, Ziqian, et al.
Pubblicazione: (2023)
Zero-Shot Voice Conversion via Content-Aware Timbre Ensemble and Conditional Flow Matching
di: Pan, Yu, et al.
Pubblicazione: (2024)
di: Pan, Yu, et al.
Pubblicazione: (2024)
The ICASSP 2026 Automatic Song Aesthetics Evaluation Challenge
di: Ma, Guobin, et al.
Pubblicazione: (2026)
di: Ma, Guobin, et al.
Pubblicazione: (2026)
Drop the beat! Freestyler for Accompaniment Conditioned Rapping Voice Generation
di: Ning, Ziqian, et al.
Pubblicazione: (2024)
di: Ning, Ziqian, et al.
Pubblicazione: (2024)
CoDiff-VC: A Codec-Assisted Diffusion Model for Zero-shot Voice Conversion
di: Li, Yuke, et al.
Pubblicazione: (2024)
di: Li, Yuke, et al.
Pubblicazione: (2024)
StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching
di: Yao, Jixun, et al.
Pubblicazione: (2024)
di: Yao, Jixun, et al.
Pubblicazione: (2024)
MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
di: Ma, Guobin, et al.
Pubblicazione: (2025)
di: Ma, Guobin, et al.
Pubblicazione: (2025)
SingVisio: Visual Analytics of Diffusion Model for Singing Voice Conversion
di: Xue, Liumeng, et al.
Pubblicazione: (2024)
di: Xue, Liumeng, et al.
Pubblicazione: (2024)
Leveraging Diverse Semantic-based Audio Pretrained Models for Singing Voice Conversion
di: Zhang, Xueyao, et al.
Pubblicazione: (2023)
di: Zhang, Xueyao, et al.
Pubblicazione: (2023)
The THU-HCSI Multi-Speaker Multi-Lingual Few-Shot Voice Cloning System for LIMMITS'24 Challenge
di: Zhou, Yixuan, et al.
Pubblicazione: (2024)
di: Zhou, Yixuan, et al.
Pubblicazione: (2024)
Semantic-Aware Interruption Detection in Spoken Dialogue Systems: Benchmark, Metric, and Model
di: Xia, Kangxiang, et al.
Pubblicazione: (2026)
di: Xia, Kangxiang, et al.
Pubblicazione: (2026)
Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling
di: Yang, Yuguang, et al.
Pubblicazione: (2024)
di: Yang, Yuguang, et al.
Pubblicazione: (2024)
SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
di: Guo, Zhao, et al.
Pubblicazione: (2025)
di: Guo, Zhao, et al.
Pubblicazione: (2025)
Summary on The Multilingual Conversational Speech Language Model Challenge: Datasets, Tasks, Baselines, and Methods
di: Mu, Bingshen, et al.
Pubblicazione: (2025)
di: Mu, Bingshen, et al.
Pubblicazione: (2025)
An Extensive Analysis of the Singing Voice Conversion Challenge 2025 Evaluation Results
di: Violeta, Lester Phillip, et al.
Pubblicazione: (2025)
di: Violeta, Lester Phillip, et al.
Pubblicazione: (2025)
ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech
di: Pan, Yu, et al.
Pubblicazione: (2025)
di: Pan, Yu, et al.
Pubblicazione: (2025)
An Initial Investigation of Neural Replay Simulator for Over-the-Air Adversarial Perturbations to Automatic Speaker Verification
di: Li, Jiaqi, et al.
Pubblicazione: (2023)
di: Li, Jiaqi, et al.
Pubblicazione: (2023)
CoMoSVC: Consistency Model-based Singing Voice Conversion
di: Lu, Yiwen, et al.
Pubblicazione: (2024)
di: Lu, Yiwen, et al.
Pubblicazione: (2024)
Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation
di: Lee, Jaejun, et al.
Pubblicazione: (2025)
di: Lee, Jaejun, et al.
Pubblicazione: (2025)
OpenVoice: Versatile Instant Voice Cloning
di: Qin, Zengyi, et al.
Pubblicazione: (2023)
di: Qin, Zengyi, et al.
Pubblicazione: (2023)
Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis
di: Zhou, Zhitong, et al.
Pubblicazione: (2025)
di: Zhou, Zhitong, et al.
Pubblicazione: (2025)
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
di: Wang, Zhichao, et al.
Pubblicazione: (2024)
di: Wang, Zhichao, et al.
Pubblicazione: (2024)
Voice Cloning: Comprehensive Survey
di: Azzuni, Hussam, et al.
Pubblicazione: (2025)
di: Azzuni, Hussam, et al.
Pubblicazione: (2025)
BridgeVoC: Revitalizing Neural Vocoder from a Restoration Perspective
di: Li, Andong, et al.
Pubblicazione: (2025)
di: Li, Andong, et al.
Pubblicazione: (2025)
Findings of the 2024 Mandarin Stuttering Event Detection and Automatic Speech Recognition Challenge
di: Xue, Hongfei, et al.
Pubblicazione: (2024)
di: Xue, Hongfei, et al.
Pubblicazione: (2024)
OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue
di: Geng, Xuelong, et al.
Pubblicazione: (2025)
di: Geng, Xuelong, et al.
Pubblicazione: (2025)
Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model
di: Zuo, Jialong, et al.
Pubblicazione: (2025)
di: Zuo, Jialong, et al.
Pubblicazione: (2025)
CoVoMix: Advancing Zero-Shot Speech Generation for Human-like Multi-talker Conversations
di: Zhang, Leying, et al.
Pubblicazione: (2024)
di: Zhang, Leying, et al.
Pubblicazione: (2024)
StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion
di: Wang, Zhichao, et al.
Pubblicazione: (2024)
di: Wang, Zhichao, et al.
Pubblicazione: (2024)
Proactive Detection of Voice Cloning with Localized Watermarking
di: Roman, Robin San, et al.
Pubblicazione: (2024)
di: Roman, Robin San, et al.
Pubblicazione: (2024)
MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech
di: Chen, Huakang, et al.
Pubblicazione: (2026)
di: Chen, Huakang, et al.
Pubblicazione: (2026)
Iterate to Differentiate: Enhancing Discriminability and Reliability in Zero-Shot TTS Evaluation
di: Shen, Shengfan, et al.
Pubblicazione: (2026)
di: Shen, Shengfan, et al.
Pubblicazione: (2026)
Voice "Cloning" is Style Transfer
di: Zhou, Kaitlyn, et al.
Pubblicazione: (2026)
di: Zhou, Kaitlyn, et al.
Pubblicazione: (2026)
Documenti analoghi
-
The NPU-HWC System for the ISCSLP 2024 Inspirational and Convincing Audio Generation Challenge
di: Guo, Dake, et al.
Pubblicazione: (2024) -
The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024
di: Zhou, Shuoyi, et al.
Pubblicazione: (2024) -
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
di: Xia, Kangxiang, et al.
Pubblicazione: (2025) -
StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding
di: Guo, Dake, et al.
Pubblicazione: (2025) -
SponTTS: modeling and transferring spontaneous style for TTS
di: Li, Hanzhao, et al.
Pubblicazione: (2023)