Synergizing Motion and Appearance: Multi-Scale Compensatory Codebooks for Talking Head Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Shuling, Hong, Fa-Ting, Huang, Xiaoshui, Xu, Dan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Online Scale Transformation for Talking Head Video Generation
by: Hong, Fa-Ting, et al.
Published: (2024)
by: Hong, Fa-Ting, et al.
Published: (2024)
Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation
by: Hong, Fa-Ting, et al.
Published: (2025)
by: Hong, Fa-Ting, et al.
Published: (2025)
Generalizable and Animatable 3D Full-Head Gaussian Avatar from a Single Image
by: Zhao, Shuling, et al.
Published: (2026)
by: Zhao, Shuling, et al.
Published: (2026)
Video Motion Graphs
by: Liu, Haiyang, et al.
Published: (2025)
by: Liu, Haiyang, et al.
Published: (2025)
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation
by: Wang, Yucheng, et al.
Published: (2025)
by: Wang, Yucheng, et al.
Published: (2025)
Total-Editing: Head Avatar with Editable Appearance, Motion, and Lighting
by: Zhao, Yizhou, et al.
Published: (2025)
by: Zhao, Yizhou, et al.
Published: (2025)
JointTuner: Appearance-Motion Adaptive Joint Training for Customized Video Generation
by: Chen, Fangda, et al.
Published: (2025)
by: Chen, Fangda, et al.
Published: (2025)
VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models
by: Chefer, Hila, et al.
Published: (2025)
by: Chefer, Hila, et al.
Published: (2025)
AtomicMotion: Learning Human Motion From Different Human Parts
by: Liu, Runzhen, et al.
Published: (2026)
by: Liu, Runzhen, et al.
Published: (2026)
THEval. Evaluation Framework for Talking Head Video Generation
by: Quignon, Nabyl, et al.
Published: (2025)
by: Quignon, Nabyl, et al.
Published: (2025)
Tora2: Motion and Appearance Customized Diffusion Transformer for Multi-Entity Video Generation
by: Zhang, Zhenghao, et al.
Published: (2025)
by: Zhang, Zhenghao, et al.
Published: (2025)
MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset
by: Sung-Bin, Kim, et al.
Published: (2024)
by: Sung-Bin, Kim, et al.
Published: (2024)
Make Your Actor Talk: Generalizable and High-Fidelity Lip Sync with Motion and Appearance Disentanglement
by: Yu, Runyi, et al.
Published: (2024)
by: Yu, Runyi, et al.
Published: (2024)
PAM: A Pose-Appearance-Motion Engine for Sim-to-Real HOI Video Generation
by: Gao, Mingju, et al.
Published: (2026)
by: Gao, Mingju, et al.
Published: (2026)
FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis
by: Wang, Mengchao, et al.
Published: (2025)
by: Wang, Mengchao, et al.
Published: (2025)
MoCoTalk: Multi-Conditional Diffusion with Adaptive Router for Controllable Talking Head Generation
by: Ye, Xinyan, et al.
Published: (2026)
by: Ye, Xinyan, et al.
Published: (2026)
AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement
by: Zhong, Zhizhou, et al.
Published: (2025)
by: Zhong, Zhizhou, et al.
Published: (2025)
Appearance Blur-driven AutoEncoder and Motion-guided Memory Module for Video Anomaly Detection
by: Lyu, Jiahao, et al.
Published: (2024)
by: Lyu, Jiahao, et al.
Published: (2024)
Identity-Preserving Video Dubbing Using Motion Warping
by: Liu, Runzhen, et al.
Published: (2025)
by: Liu, Runzhen, et al.
Published: (2025)
TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis
by: Chen, Shunian, et al.
Published: (2025)
by: Chen, Shunian, et al.
Published: (2025)
Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation
by: Zhai, Yuanhao, et al.
Published: (2024)
by: Zhai, Yuanhao, et al.
Published: (2024)
CustomTTT: Motion and Appearance Customized Video Generation via Test-Time Training
by: Bi, Xiuli, et al.
Published: (2024)
by: Bi, Xiuli, et al.
Published: (2024)
MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation
by: Yang, Kaixing, et al.
Published: (2025)
by: Yang, Kaixing, et al.
Published: (2025)
MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
FD2Talk: Towards Generalized Talking Head Generation with Facial Decoupled Diffusion Model
by: Yao, Ziyu, et al.
Published: (2024)
by: Yao, Ziyu, et al.
Published: (2024)
COMOGen: A Controllable Text-to-3D Multi-object Generation Framework
by: Sun, Shaorong, et al.
Published: (2024)
by: Sun, Shaorong, et al.
Published: (2024)
TalkingHeadBench: A Multi-Modal Benchmark & Analysis of Talking-Head DeepFake Detection
by: Xiong, Xinqi, et al.
Published: (2025)
by: Xiong, Xinqi, et al.
Published: (2025)
MoDiTalker: Motion-Disentangled Diffusion Model for High-Fidelity Talking Head Generation
by: Kim, Seyeon, et al.
Published: (2024)
by: Kim, Seyeon, et al.
Published: (2024)
IF-MDM: Implicit Face Motion Diffusion Model for High-Fidelity Realtime Talking Head Generation
by: Yang, Sejong, et al.
Published: (2024)
by: Yang, Sejong, et al.
Published: (2024)
DAM-VSR: Disentanglement of Appearance and Motion for Video Super-Resolution
by: Kong, Zhe, et al.
Published: (2025)
by: Kong, Zhe, et al.
Published: (2025)
SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis
by: Peng, Ziqiao, et al.
Published: (2023)
by: Peng, Ziqiao, et al.
Published: (2023)
FlexAM: Flexible Appearance-Motion Decomposition for Versatile Video Generation Control
by: Sheng, Mingzhi, et al.
Published: (2026)
by: Sheng, Mingzhi, et al.
Published: (2026)
Talking Head Generation via AU-Guided Landmark Prediction
by: Chang, Shao-Yu, et al.
Published: (2025)
by: Chang, Shao-Yu, et al.
Published: (2025)
Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router
by: Huang, Yubo, et al.
Published: (2025)
by: Huang, Yubo, et al.
Published: (2025)
Avatar Fingerprinting for Authorized Use of Synthetic Talking-Head Videos
by: Prashnani, Ekta, et al.
Published: (2023)
by: Prashnani, Ekta, et al.
Published: (2023)
UniAvatar: Taming Lifelike Audio-Driven Talking Head Generation with Comprehensive Motion and Lighting Control
by: Sun, Wenzhang, et al.
Published: (2024)
by: Sun, Wenzhang, et al.
Published: (2024)
Video Anomaly Detection with Motion and Appearance Guided Patch Diffusion Model
by: Zhou, Hang, et al.
Published: (2024)
by: Zhou, Hang, et al.
Published: (2024)
TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking Styles
by: Ma, Yifeng, et al.
Published: (2023)
by: Ma, Yifeng, et al.
Published: (2023)
HM-Talker: Hybrid Motion Modeling for High-Fidelity Talking Head Synthesis
by: Liu, Shiyu, et al.
Published: (2025)
by: Liu, Shiyu, et al.
Published: (2025)
RSATalker: Realistic Socially-Aware Talking Head Generation for Multi-Turn Conversation
by: Chen, Peng, et al.
Published: (2026)
by: Chen, Peng, et al.
Published: (2026)
Similar Items
-
Learning Online Scale Transformation for Talking Head Video Generation
by: Hong, Fa-Ting, et al.
Published: (2024) -
Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation
by: Hong, Fa-Ting, et al.
Published: (2025) -
Generalizable and Animatable 3D Full-Head Gaussian Avatar from a Single Image
by: Zhao, Shuling, et al.
Published: (2026) -
Video Motion Graphs
by: Liu, Haiyang, et al.
Published: (2025) -
MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation
by: Wang, Yucheng, et al.
Published: (2025)