MimicTalk: Mimicking a personalized and expressive 3D talking face in minutes
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Zhenhui, Zhong, Tianyun, Ren, Yi, Jiang, Ziyue, Huang, Jiawei, Huang, Rongjie, Liu, Jinglin, He, Jinzheng, Zhang, Chen, Wang, Zehan, Chen, Xize, Yin, Xiang, Zhao, Zhou |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis
by: Ye, Zhenhui, et al.
Published: (2024)
by: Ye, Zhenhui, et al.
Published: (2024)
MulliVC: Multi-lingual Voice Conversion With Cycle Consistency
by: Huang, Jiawei, et al.
Published: (2024)
by: Huang, Jiawei, et al.
Published: (2024)
Unleashing the Power of Natural Audio Featuring Multiple Sound Sources
by: Cheng, Xize, et al.
Published: (2025)
by: Cheng, Xize, et al.
Published: (2025)
FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion
by: Wang, Zehan, et al.
Published: (2024)
by: Wang, Zehan, et al.
Published: (2024)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
by: Jiang, Ziyue, et al.
Published: (2023)
by: Jiang, Ziyue, et al.
Published: (2023)
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
by: Wang, Zehan, et al.
Published: (2025)
by: Wang, Zehan, et al.
Published: (2025)
OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
by: Wang, Zehan, et al.
Published: (2024)
by: Wang, Zehan, et al.
Published: (2024)
Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
by: Huang, Haifeng, et al.
Published: (2023)
by: Huang, Haifeng, et al.
Published: (2023)
TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
Generate Your Talking Avatar from Video Reference
by: Guo, Zujin, et al.
Published: (2026)
by: Guo, Zujin, et al.
Published: (2026)
Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching
by: Wang, Yongqi, et al.
Published: (2024)
by: Wang, Yongqi, et al.
Published: (2024)
ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
by: Ji, Shengpeng, et al.
Published: (2024)
by: Ji, Shengpeng, et al.
Published: (2024)
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
by: Cheng, Xize, et al.
Published: (2024)
by: Cheng, Xize, et al.
Published: (2024)
WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
by: Ji, Shengpeng, et al.
Published: (2024)
by: Ji, Shengpeng, et al.
Published: (2024)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
by: Liu, Huadai, et al.
Published: (2023)
by: Liu, Huadai, et al.
Published: (2023)
On the exponent of distribution for convolutions of $\mathrm{GL(2)}$ coefficients to smooth moduli
by: Yin, Rongjie
Published: (2025)
by: Yin, Rongjie
Published: (2025)
HybridMimic: Hybrid RL-Centroidal Control for Humanoid Motion Mimicking
by: Tay, Ludwig Chee-Ying, et al.
Published: (2026)
by: Tay, Ludwig Chee-Ying, et al.
Published: (2026)
Stay, leave late, leave early, return, or move onward? Interprovincial migration decisions of older adults in China, 2000–2005 and 2010–2015
by: Cuiying Huang, et al.
Published: (2024)
by: Cuiying Huang, et al.
Published: (2024)
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
by: Cheng, Xize, et al.
Published: (2025)
by: Cheng, Xize, et al.
Published: (2025)
Text-to-Song: Towards Controllable Music Generation Incorporating Vocals and Accompaniment
by: Hong, Zhiqing, et al.
Published: (2024)
by: Hong, Zhiqing, et al.
Published: (2024)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
by: Liu, Huadai, et al.
Published: (2024)
by: Liu, Huadai, et al.
Published: (2024)
Superior and Pragmatic Talking Face Generation with Teacher-Student Framework
by: Liang, Chao, et al.
Published: (2024)
by: Liang, Chao, et al.
Published: (2024)
Unsupervised Solution Operator Learning for Mean-Field Games via Sampling-Invariant Parametrizations
by: Huang, Han, et al.
Published: (2024)
by: Huang, Han, et al.
Published: (2024)
Talking to oneself in CMC: a study of self replies in Wikipedia talk pages
by: Tanguy, Ludovic, et al.
Published: (2024)
by: Tanguy, Ludovic, et al.
Published: (2024)
Joint Inference of Trajectory and Obstacle in Mean-Field Games via Bilevel Optimization
by: Huang, Han, et al.
Published: (2025)
by: Huang, Han, et al.
Published: (2025)
On the exponent of distribution for convolutions of $\operatorname{GL}(2)$ coefficients to smooth moduli
by: Yin, Rongjie, et al.
Published: (2026)
by: Yin, Rongjie, et al.
Published: (2026)
SegTalker: Segmentation-based Talking Face Generation with Mask-guided Local Editing
by: Xiong, Lingyu, et al.
Published: (2024)
by: Xiong, Lingyu, et al.
Published: (2024)
Biomimetic Vascularized iPSC‐Hepatocyte Spheroids for Liver Regeneration
by: Jinglin Wang, et al.
Published: (2024)
by: Jinglin Wang, et al.
Published: (2024)
MerLean-Prover: A Recursive Looping Harness for Lean 4 Theorem Proving
by: Li, Jinzheng, et al.
Published: (2026)
by: Li, Jinzheng, et al.
Published: (2026)
MerLean: An Agentic Framework for Autoformalization in Quantum Computation
by: Ren, Yuanjie, et al.
Published: (2026)
by: Ren, Yuanjie, et al.
Published: (2026)
Bacteria Flagella‐Mimicking Polymer Multilayer Magnetic Microrobots
by: Liang Lu, et al.
Published: (2025)
by: Liang Lu, et al.
Published: (2025)
Application of hypothetical strategies in acute pain
by: Jinglin Zhong, et al.
Published: (2024)
by: Jinglin Zhong, et al.
Published: (2024)
4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos
by: Zhong, Shanshan, et al.
Published: (2025)
by: Zhong, Shanshan, et al.
Published: (2025)
HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation
by: Gan, Qijun, et al.
Published: (2025)
by: Gan, Qijun, et al.
Published: (2025)
VDE: Training-Free Accelerating Rectified Flow Model via Velocity Decomposition and Estimation
by: Tan, Junwen, et al.
Published: (2026)
by: Tan, Junwen, et al.
Published: (2026)
Emotion recognition in talking-face videos using persistent entropy and neural networks
by: Paluzo-Hidalgo, Eduardo, et al.
Published: (2021)
by: Paluzo-Hidalgo, Eduardo, et al.
Published: (2021)
WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
by: Ji, Shengpeng, et al.
Published: (2025)
by: Ji, Shengpeng, et al.
Published: (2025)
Mimic Intent, Not Just Trajectories
by: Huang, Renming, et al.
Published: (2026)
by: Huang, Renming, et al.
Published: (2026)
Talking the talk and walking the walk: Aquaculture, fish, and fisheries will continue to support the blue revolution and beyond
by: Christyn Bailey, et al.
Published: (2024)
by: Christyn Bailey, et al.
Published: (2024)
TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis
by: Chen, Shunian, et al.
Published: (2025)
by: Chen, Shunian, et al.
Published: (2025)
Similar Items
-
Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis
by: Ye, Zhenhui, et al.
Published: (2024) -
MulliVC: Multi-lingual Voice Conversion With Cycle Consistency
by: Huang, Jiawei, et al.
Published: (2024) -
Unleashing the Power of Natural Audio Featuring Multiple Sound Sources
by: Cheng, Xize, et al.
Published: (2025) -
FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion
by: Wang, Zehan, et al.
Published: (2024) -
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
by: Jiang, Ziyue, et al.
Published: (2023)