Multi-identity Human Image Animation with Structural Video Diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zhenzhi, Li, Yixuan, Zeng, Yanhong, Guo, Yuwei, Lin, Dahua, Xue, Tianfan, Dai, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HumanVid: Demystifying Training Data for Camera-controllable Human Image Animation
by: Wang, Zhenzhi, et al.
Published: (2024)
by: Wang, Zhenzhi, et al.
Published: (2024)
InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions
by: Wang, Zhenzhi, et al.
Published: (2025)
by: Wang, Zhenzhi, et al.
Published: (2025)
InterControl: Zero-shot Human Interaction Generation by Controlling Every Joint
by: Wang, Zhenzhi, et al.
Published: (2023)
by: Wang, Zhenzhi, et al.
Published: (2023)
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
by: Wang, Zhenzhi, et al.
Published: (2025)
by: Wang, Zhenzhi, et al.
Published: (2025)
PIA: Your Personalized Image Animator via Plug-and-Play Modules in Text-to-Image Models
by: Zhang, Yiming, et al.
Published: (2023)
by: Zhang, Yiming, et al.
Published: (2023)
StableAnimator: High-Quality Identity-Preserving Human Image Animation
by: Tu, Shuyuan, et al.
Published: (2024)
by: Tu, Shuyuan, et al.
Published: (2024)
EvAnimate: Event-conditioned Image-to-Video Generation for Human Animation
by: Qu, Qiang, et al.
Published: (2025)
by: Qu, Qiang, et al.
Published: (2025)
StableAnimator++: Overcoming Pose Misalignment and Face Distortion for Human Image Animation
by: Tu, Shuyuan, et al.
Published: (2025)
by: Tu, Shuyuan, et al.
Published: (2025)
AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
by: Guo, Yuwei, et al.
Published: (2023)
by: Guo, Yuwei, et al.
Published: (2023)
MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model
by: Niu, Muyao, et al.
Published: (2024)
by: Niu, Muyao, et al.
Published: (2024)
TCAN: Animating Human Images with Temporally Consistent Pose Guidance using Diffusion Models
by: Kim, Jeongho, et al.
Published: (2024)
by: Kim, Jeongho, et al.
Published: (2024)
iHuman: Instant Animatable Digital Humans From Monocular Videos
by: Paudel, Pramish, et al.
Published: (2024)
by: Paudel, Pramish, et al.
Published: (2024)
Implicit Preference Alignment for Human Image Animation
by: Wang, Yuanzhi, et al.
Published: (2026)
by: Wang, Yuanzhi, et al.
Published: (2026)
Animate Any Character in Any World
by: Wang, Yitong, et al.
Published: (2025)
by: Wang, Yitong, et al.
Published: (2025)
AnimateDiff-Lightning: Cross-Model Diffusion Distillation
by: Lin, Shanchuan, et al.
Published: (2024)
by: Lin, Shanchuan, et al.
Published: (2024)
LHM: Large Animatable Human Reconstruction Model from a Single Image in Seconds
by: Qiu, Lingteng, et al.
Published: (2025)
by: Qiu, Lingteng, et al.
Published: (2025)
MVP4D: Multi-View Portrait Video Diffusion for Animatable 4D Avatars
by: Taubner, Felix, et al.
Published: (2025)
by: Taubner, Felix, et al.
Published: (2025)
RelightVid: Temporal-Consistent Diffusion Model for Video Relighting
by: Fang, Ye, et al.
Published: (2025)
by: Fang, Ye, et al.
Published: (2025)
EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration
by: Li, Wuyang, et al.
Published: (2026)
by: Li, Wuyang, et al.
Published: (2026)
CubeComposer: Spatio-Temporal Autoregressive 4K 360° Video Generation from Perspective Video
by: Li, Lingen, et al.
Published: (2026)
by: Li, Lingen, et al.
Published: (2026)
ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
by: Zhang, Mengchen, et al.
Published: (2025)
by: Zhang, Mengchen, et al.
Published: (2025)
Temporal Aware Pruning for Efficient Diffusion-based Video Generation
by: Li, Sheng, et al.
Published: (2026)
by: Li, Sheng, et al.
Published: (2026)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
X-UniMotion: Animating Human Images with Expressive, Unified and Identity-Agnostic Motion Latents
by: Song, Guoxian, et al.
Published: (2025)
by: Song, Guoxian, et al.
Published: (2025)
DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid Guidance
by: Luo, Yuxuan, et al.
Published: (2025)
by: Luo, Yuxuan, et al.
Published: (2025)
Animating the Past: Reconstruct Trilobite via Video Generation
by: Wu, Xiaoran, et al.
Published: (2024)
by: Wu, Xiaoran, et al.
Published: (2024)
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
by: Gan, Qijun, et al.
Published: (2025)
by: Gan, Qijun, et al.
Published: (2025)
Composing Concepts from Images and Videos via Concept-prompt Binding
by: Kong, Xianghao, et al.
Published: (2025)
by: Kong, Xianghao, et al.
Published: (2025)
Versatile Multimodal Controls for Expressive Talking Human Animation
by: Qin, Zheng, et al.
Published: (2025)
by: Qin, Zheng, et al.
Published: (2025)
VideoPanda: Video Panoramic Diffusion with Multi-view Attention
by: Xie, Kevin, et al.
Published: (2025)
by: Xie, Kevin, et al.
Published: (2025)
Dormant: Defending against Pose-driven Human Image Animation
by: Zhou, Jiachen, et al.
Published: (2024)
by: Zhou, Jiachen, et al.
Published: (2024)
MAVIN: Multi-Action Video Generation with Diffusion Models via Transition Video Infilling
by: Zhang, Bowen, et al.
Published: (2024)
by: Zhang, Bowen, et al.
Published: (2024)
Embedded Representation Learning Network for Animating Styled Video Portrait
by: Wang, Tianyong, et al.
Published: (2024)
by: Wang, Tianyong, et al.
Published: (2024)
Training-free Dense-Aligned Diffusion Guidance for Modular Conditional Image Synthesis
by: Wang, Zixuan, et al.
Published: (2025)
by: Wang, Zixuan, et al.
Published: (2025)
Interactive Character Control with Auto-Regressive Motion Diffusion Models
by: Shi, Yi, et al.
Published: (2023)
by: Shi, Yi, et al.
Published: (2023)
Every Image Listens, Every Image Dances: Music-Driven Image Animation
by: Dong, Zhikang, et al.
Published: (2025)
by: Dong, Zhikang, et al.
Published: (2025)
FADE: Frequency-Aware Diffusion Model Factorization for Video Editing
by: Zhu, Yixuan, et al.
Published: (2025)
by: Zhu, Yixuan, et al.
Published: (2025)
AdvI2I: Adversarial Image Attack on Image-to-Image Diffusion models
by: Zeng, Yaopei, et al.
Published: (2024)
by: Zeng, Yaopei, et al.
Published: (2024)
LoopAnimate: Loopable Salient Object Animation
by: Wang, Fanyi, et al.
Published: (2024)
by: Wang, Fanyi, et al.
Published: (2024)
Demystifying Video Reasoning
by: Wang, Ruisi, et al.
Published: (2026)
by: Wang, Ruisi, et al.
Published: (2026)
Similar Items
-
HumanVid: Demystifying Training Data for Camera-controllable Human Image Animation
by: Wang, Zhenzhi, et al.
Published: (2024) -
InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions
by: Wang, Zhenzhi, et al.
Published: (2025) -
InterControl: Zero-shot Human Interaction Generation by Controlling Every Joint
by: Wang, Zhenzhi, et al.
Published: (2023) -
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
by: Wang, Zhenzhi, et al.
Published: (2025) -
PIA: Your Personalized Image Animator via Plug-and-Play Modules in Text-to-Image Models
by: Zhang, Yiming, et al.
Published: (2023)