Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Chen, Sijing, Feng, Yuan, He, Laipeng, He, Tianwei, He, Wendi, Hu, Yanni, Lin, Bin, Lin, Yiting, Pan, Yu, Tan, Pengfei, Tian, Chengwei, Wang, Chen, Wang, Zhicheng, Xie, Ruoye, Yao, Jixun, Yan, Quanlei, Yang, Yuguang, Ye, Jianhao, Yin, Jingjing, Yu, Yanzhen, Zhang, Huimin, Zhang, Xiang, Zhao, Guangcheng, Zhou, Hongbin, Zou, Pengpeng |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Takin-ADA: Emotion Controllable Audio-Driven Animation with Canonical and Landmark Loss Optimization
par: Lin, Bin, et autres
Publié: (2024)
par: Lin, Bin, et autres
Publié: (2024)
Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling
par: Yang, Yuguang, et autres
Publié: (2024)
par: Yang, Yuguang, et autres
Publié: (2024)
ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech
par: Pan, Yu, et autres
Publié: (2025)
par: Pan, Yu, et autres
Publié: (2025)
Fine-grained Preference Optimization Improves Zero-shot Text-to-Speech
par: Yao, Jixun, et autres
Publié: (2025)
par: Yao, Jixun, et autres
Publié: (2025)
PSCodec: A Series of High-Fidelity Low-bitrate Neural Speech Codecs Leveraging Prompt Encoders
par: Pan, Yu, et autres
Publié: (2024)
par: Pan, Yu, et autres
Publié: (2024)
VG-Mapping: Variation-aware Density Control for Online 3D Gaussian Mapping in Semi-static Scenes
par: He, Yicheng, et autres
Publié: (2025)
par: He, Yicheng, et autres
Publié: (2025)
Optimizing NeRF-based SLAM with Trajectory Smoothness Constraints
par: He, Yicheng, et autres
Publié: (2024)
par: He, Yicheng, et autres
Publié: (2024)
PISR: Polarimetric Neural Implicit Surface Reconstruction for Textureless and Specular Objects
par: Chen, Guangcheng, et autres
Publié: (2024)
par: Chen, Guangcheng, et autres
Publié: (2024)
Isolation and Characterization of Mesenchymal Stem Cells Derived From the Bone Marrow of Takin ( Budorcas taxicolor )
par: Wei‐Qiang Luo, et autres
Publié: (2025)
par: Wei‐Qiang Luo, et autres
Publié: (2025)
StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching
par: Yao, Jixun, et autres
Publié: (2024)
par: Yao, Jixun, et autres
Publié: (2024)
PolygMap: A Perceptive Locomotion Framework for Humanoid Robot Stair Climbing
par: Li, Bingquan, et autres
Publié: (2025)
par: Li, Bingquan, et autres
Publié: (2025)
Achieving Ultrahigh Voltage Over 100 V and Remarkable Freshwater Harvesting Based on Thermodiffusion Enhanced Hydrovoltaic Generator
par: Yu Chen, et autres
Publié: (2024)
par: Yu Chen, et autres
Publié: (2024)
Achieving Ultrahigh Voltage Over 100 V and Remarkable Freshwater Harvesting Based on Thermodiffusion Enhanced Hydrovoltaic Generator (Adv. Energy Mater. 24/2024)
par: Yu Chen, et autres
Publié: (2024)
par: Yu Chen, et autres
Publié: (2024)
Positive mass theorems on singular spaces and some applications
par: He, Shihang, et autres
Publié: (2025)
par: He, Shihang, et autres
Publié: (2025)
Foliation of area-minimizing hypersurfaces in asymptotically flat manifolds of higher dimension
par: He, Shihang, et autres
Publié: (2026)
par: He, Shihang, et autres
Publié: (2026)
Singularity removal rigidity theorems for minimal hypersurfaces in manifolds with nonnegative scalar curvature
par: He, Shihang, et autres
Publié: (2026)
par: He, Shihang, et autres
Publié: (2026)
Foliation of area minimizing hypersurfaces in asymptotically flat manifolds and Schoen's conjecture
par: He, Shihang, et autres
Publié: (2024)
par: He, Shihang, et autres
Publié: (2024)
MSAC: Multiple Speech Attribute Control Method for Reliable Speech Emotion Recognition
par: Pan, Yu, et autres
Publié: (2023)
par: Pan, Yu, et autres
Publié: (2023)
Zero-Shot Voice Conversion via Content-Aware Timbre Ensemble and Conditional Flow Matching
par: Pan, Yu, et autres
Publié: (2024)
par: Pan, Yu, et autres
Publié: (2024)
China’s Two Child Policy: Can a Name Change Mean a Game Change?
par: Chen Guangcheng
Publié: (2016)
par: Chen Guangcheng
Publié: (2016)
Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy
par: Ma, Linhan, et autres
Publié: (2024)
par: Ma, Linhan, et autres
Publié: (2024)
Morphological Study on the Coat Hair in Golden Takin
par: Yutaka Kawahara, et autres
Publié: (2025)
par: Yutaka Kawahara, et autres
Publié: (2025)
Enhancing Glass Surface Reconstruction via Depth Prior for Robot Navigation
par: Zheng, Jiamin, et autres
Publié: (2026)
par: Zheng, Jiamin, et autres
Publié: (2026)
Hierarchical Visual Relocalization with Nearest View Synthesis from Feature Gaussian Splatting
par: Tao, Huaqi, et autres
Publié: (2026)
par: Tao, Huaqi, et autres
Publié: (2026)
Audiobook-CC: Controllable Long-context Speech Generation for Multicast Audiobook
par: Liu, Min, et autres
Publié: (2025)
par: Liu, Min, et autres
Publié: (2025)
Catch the Butterfly: Peeking into the Terms and Conflicts among SPDX Licenses
par: Liu, Tao, et autres
Publié: (2024)
par: Liu, Tao, et autres
Publié: (2024)
S2ST-Omni: Hierarchical Language-Aware SpeechLLM Adaptation for Multilingual Speech-to-Speech Translation
par: Pan, Yu, et autres
Publié: (2025)
par: Pan, Yu, et autres
Publié: (2025)
Psychological Mechanisms of Generative AI Discontinuance Intention among Chinese K-12 Teachers
par: Du, Yiran, et autres
Publié: (2026)
par: Du, Yiran, et autres
Publié: (2026)
Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning
par: Lin, Juekai, et autres
Publié: (2026)
par: Lin, Juekai, et autres
Publié: (2026)
Quantum Algorithms for Projection-Free Sparse Convex Optimization
par: He, Jianhao, et autres
Publié: (2025)
par: He, Jianhao, et autres
Publié: (2025)
Disorder effects on the Topological Superconductor with Hubbard Interactions
par: Deng, Yiting, et autres
Publié: (2023)
par: Deng, Yiting, et autres
Publié: (2023)
The Quadrupole Moment of Higher-Order Topological Insulator at Finite temperature
par: Deng, Yiting, et autres
Publié: (2025)
par: Deng, Yiting, et autres
Publié: (2025)
AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection
par: Zhou, Qihang, et autres
Publié: (2023)
par: Zhou, Qihang, et autres
Publié: (2023)
EyecareGPT: Boosting Comprehensive Ophthalmology Understanding with Tailored Dataset, Benchmark and Model
par: Li, Sijing, et autres
Publié: (2025)
par: Li, Sijing, et autres
Publié: (2025)
Strengthen or Weaken: Evolutionary Directions of Cross‐Feeding After Formation
par: Laipeng Luo, et autres
Publié: (2025)
par: Laipeng Luo, et autres
Publié: (2025)
Effect of Cerium Addition on Inclusion Evolution and Mechanical Performance of Wind Power Steel
par: Ziyan Cheng, et autres
Publié: (2026)
par: Ziyan Cheng, et autres
Publié: (2026)
Learning Fair Domain Adaptation with Virtual Label Distribution
par: Zhang, Yuguang, et autres
Publié: (2026)
par: Zhang, Yuguang, et autres
Publié: (2026)
Vacancy Engineering in the First Coordination Shell of Single‐Atom Catalysts for Enhanced Hydrogen and Oxygen Evolution Reactions
par: Xinqi Chen, et autres
Publié: (2025)
par: Xinqi Chen, et autres
Publié: (2025)
Siamese Transformer Networks for Few-shot Image Classification
par: Jiang, Weihao, et autres
Publié: (2024)
par: Jiang, Weihao, et autres
Publié: (2024)
Why Learners Drift In and Out: Examining Intermittent Discontinuance in AI-Mediated Informal Digital English Learning (AI-IDLE) Using SEM and fsQCA
par: Du, Yiran, et autres
Publié: (2026)
par: Du, Yiran, et autres
Publié: (2026)
Documents similaires
-
Takin-ADA: Emotion Controllable Audio-Driven Animation with Canonical and Landmark Loss Optimization
par: Lin, Bin, et autres
Publié: (2024) -
Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling
par: Yang, Yuguang, et autres
Publié: (2024) -
ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech
par: Pan, Yu, et autres
Publié: (2025) -
Fine-grained Preference Optimization Improves Zero-shot Text-to-Speech
par: Yao, Jixun, et autres
Publié: (2025) -
PSCodec: A Series of High-Fidelity Low-bitrate Neural Speech Codecs Leveraging Prompt Encoders
par: Pan, Yu, et autres
Publié: (2024)