The FruitShell French synthesis system at the Blizzard 2023 Challenge
Fuente:
arXiv
Salvato in:
| Autori principali: | Qi, Xin, Wang, Xiaopeng, Wang, Zhiyong, Liu, Wang, Ding, Mingming, Shi, Shuchen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A multi-speaker multi-lingual voice cloning system based on vits2 for limmits 2024 challenge
di: Wang, Xiaopeng, et al.
Pubblicazione: (2024)
di: Wang, Xiaopeng, et al.
Pubblicazione: (2024)
Genuine-Focused Learning using Mask AutoEncoder for Generalized Fake Audio Detection
di: Wang, Xiaopeng, et al.
Pubblicazione: (2024)
di: Wang, Xiaopeng, et al.
Pubblicazione: (2024)
Generalized Fake Audio Detection via Deep Stable Learning
di: Wang, Zhiyong, et al.
Pubblicazione: (2024)
di: Wang, Zhiyong, et al.
Pubblicazione: (2024)
PPPR: Portable Plug-in Prompt Refiner for Text to Audio Generation
di: Shi, Shuchen, et al.
Pubblicazione: (2024)
di: Shi, Shuchen, et al.
Pubblicazione: (2024)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
di: Fu, Ruibo, et al.
Pubblicazione: (2024)
di: Fu, Ruibo, et al.
Pubblicazione: (2024)
EELE: Exploring Efficient and Extensible LoRA Integration in Emotional Text-to-Speech
di: Qi, Xin, et al.
Pubblicazione: (2024)
di: Qi, Xin, et al.
Pubblicazione: (2024)
Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
di: Wang, Zhiyong, et al.
Pubblicazione: (2024)
di: Wang, Zhiyong, et al.
Pubblicazione: (2024)
Codecfake: An Initial Dataset for Detecting LLM-based Deepfake Audio
di: Lu, Yi, et al.
Pubblicazione: (2024)
di: Lu, Yi, et al.
Pubblicazione: (2024)
Findings of the 2023 ML-SUPERB Challenge: Pre-Training and Evaluation over More Languages and Beyond
di: Shi, Jiatong, et al.
Pubblicazione: (2023)
di: Shi, Jiatong, et al.
Pubblicazione: (2023)
Bridging Language Gaps in Audio-Text Retrieval
di: Yan, Zhiyong, et al.
Pubblicazione: (2024)
di: Yan, Zhiyong, et al.
Pubblicazione: (2024)
An audio-quality-based multi-strategy approach for target speaker extraction in the MISP 2023 Challenge
di: Han, Runduo, et al.
Pubblicazione: (2024)
di: Han, Runduo, et al.
Pubblicazione: (2024)
A Noval Feature via Color Quantisation for Fake Audio Detection
di: Wang, Zhiyong, et al.
Pubblicazione: (2024)
di: Wang, Zhiyong, et al.
Pubblicazione: (2024)
Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding
di: Liu, Jizhong, et al.
Pubblicazione: (2024)
di: Liu, Jizhong, et al.
Pubblicazione: (2024)
System Description for the Displace Speaker Diarization Challenge 2023
di: Aliyev, Ali
Pubblicazione: (2024)
di: Aliyev, Ali
Pubblicazione: (2024)
DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
di: Qi, Xin, et al.
Pubblicazione: (2024)
di: Qi, Xin, et al.
Pubblicazione: (2024)
Pitch-Aware RNN-T for Mandarin Chinese Mispronunciation Detection and Diagnosis
di: Wang, Xintong, et al.
Pubblicazione: (2024)
di: Wang, Xintong, et al.
Pubblicazione: (2024)
The THUEE System Description for the IARPA OpenASR21 Challenge
di: Zhao, Jing, et al.
Pubblicazione: (2022)
di: Zhao, Jing, et al.
Pubblicazione: (2022)
Temporal Variability and Multi-Viewed Self-Supervised Representations to Tackle the ASVspoof5 Deepfake Challenge
di: Xie, Yuankun, et al.
Pubblicazione: (2024)
di: Xie, Yuankun, et al.
Pubblicazione: (2024)
SAFE-QAQ: End-to-End Slow-Thinking Audio-Text Fraud Detection via Reinforcement Learning
di: Wang, Peidong, et al.
Pubblicazione: (2026)
di: Wang, Peidong, et al.
Pubblicazione: (2026)
GLAP: General contrastive audio-text pretraining across domains and languages
di: Dinkel, Heinrich, et al.
Pubblicazione: (2025)
di: Dinkel, Heinrich, et al.
Pubblicazione: (2025)
The Third VoicePrivacy Challenge: Preserving Emotional Expressiveness and Linguistic Content in Voice Anonymization
di: Tomashenko, Natalia, et al.
Pubblicazione: (2026)
di: Tomashenko, Natalia, et al.
Pubblicazione: (2026)
Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data
di: Xie, Jingran, et al.
Pubblicazione: (2025)
di: Xie, Jingran, et al.
Pubblicazione: (2025)
Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving
di: Xie, Jingran, et al.
Pubblicazione: (2025)
di: Xie, Jingran, et al.
Pubblicazione: (2025)
CAFE A Novel Code switching Dataset for Algerian Dialect French and English
di: Lachemat, Houssam Eddine-Othman, et al.
Pubblicazione: (2024)
di: Lachemat, Houssam Eddine-Othman, et al.
Pubblicazione: (2024)
The Interspeech 2026 Audio Encoder Capability Challenge for Large Audio Language Models
di: Dinkel, Heinrich, et al.
Pubblicazione: (2026)
di: Dinkel, Heinrich, et al.
Pubblicazione: (2026)
RPRA-ADD: Forgery Trace Enhancement-Driven Audio Deepfake Detection
di: Fu, Ruibo, et al.
Pubblicazione: (2025)
di: Fu, Ruibo, et al.
Pubblicazione: (2025)
Evolution of Voices in French Audiovisual Media Across Genders and Age in a Diachronic Perspective
di: Rilliard, Albert, et al.
Pubblicazione: (2024)
di: Rilliard, Albert, et al.
Pubblicazione: (2024)
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview
di: Liu, Heyang, et al.
Pubblicazione: (2024)
di: Liu, Heyang, et al.
Pubblicazione: (2024)
MM-KWS: Multi-modal Prompts for Multilingual User-defined Keyword Spotting
di: Ai, Zhiqi, et al.
Pubblicazione: (2024)
di: Ai, Zhiqi, et al.
Pubblicazione: (2024)
Advances in Speech Separation: Techniques, Challenges, and Future Trends
di: Li, Kai, et al.
Pubblicazione: (2025)
di: Li, Kai, et al.
Pubblicazione: (2025)
Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding
di: Zhao, Jiahui, et al.
Pubblicazione: (2024)
di: Zhao, Jiahui, et al.
Pubblicazione: (2024)
The NTNU System at the S&I Challenge 2025 SLA Open Track
di: Lin, Hong-Yun, et al.
Pubblicazione: (2025)
di: Lin, Hong-Yun, et al.
Pubblicazione: (2025)
A Benchmark for Multi-speaker Anonymization
di: Miao, Xiaoxiao, et al.
Pubblicazione: (2024)
di: Miao, Xiaoxiao, et al.
Pubblicazione: (2024)
The ICME 2025 Audio Encoder Capability Challenge
di: Zhang, Junbo, et al.
Pubblicazione: (2025)
di: Zhang, Junbo, et al.
Pubblicazione: (2025)
StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion
di: Li, Fengjin, et al.
Pubblicazione: (2025)
di: Li, Fengjin, et al.
Pubblicazione: (2025)
Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language
di: Abu, Turi, et al.
Pubblicazione: (2025)
di: Abu, Turi, et al.
Pubblicazione: (2025)
Decoding Linguistic Representations of Human Brain
di: Wang, Yu, et al.
Pubblicazione: (2024)
di: Wang, Yu, et al.
Pubblicazione: (2024)
Continual Learning Optimizations for Auto-regressive Decoder of Multilingual ASR systems
di: Kwok, Chin Yuen, et al.
Pubblicazione: (2024)
di: Kwok, Chin Yuen, et al.
Pubblicazione: (2024)
Streaming Audio Transformers for Online Audio Tagging
di: Dinkel, Heinrich, et al.
Pubblicazione: (2023)
di: Dinkel, Heinrich, et al.
Pubblicazione: (2023)
Scaling up masked audio encoder learning for general audio classification
di: Dinkel, Heinrich, et al.
Pubblicazione: (2024)
di: Dinkel, Heinrich, et al.
Pubblicazione: (2024)
Documenti analoghi
-
A multi-speaker multi-lingual voice cloning system based on vits2 for limmits 2024 challenge
di: Wang, Xiaopeng, et al.
Pubblicazione: (2024) -
Genuine-Focused Learning using Mask AutoEncoder for Generalized Fake Audio Detection
di: Wang, Xiaopeng, et al.
Pubblicazione: (2024) -
Generalized Fake Audio Detection via Deep Stable Learning
di: Wang, Zhiyong, et al.
Pubblicazione: (2024) -
PPPR: Portable Plug-in Prompt Refiner for Text to Audio Generation
di: Shi, Shuchen, et al.
Pubblicazione: (2024) -
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
di: Fu, Ruibo, et al.
Pubblicazione: (2024)