VividPose: Advancing Stable Video Diffusion for Realistic Human Image Animation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Qilin, Jiang, Zhengkai, Xu, Chengming, Zhang, Jiangning, Wang, Yabiao, Zhang, Xinyi, Cao, Yun, Cao, Weijian, Wang, Chengjie, Fu, Yanwei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SwiftVideo: A Unified Framework for Few-Step Video Generation through Trajectory-Distribution Alignment
by: Sun, Yanxiao, et al.
Published: (2025)
by: Sun, Yanxiao, et al.
Published: (2025)
Collaborative Face Experts Fusion in Video Generation: Boosting Identity Consistency Across Large Face Poses
by: Wang, Yuji, et al.
Published: (2025)
by: Wang, Yuji, et al.
Published: (2025)
ArtWeaver: Advanced Dynamic Style Integration via Diffusion Model
by: Xu, Chengming, et al.
Published: (2024)
by: Xu, Chengming, et al.
Published: (2024)
DiffFAE: Advancing High-fidelity One-shot Facial Appearance Editing with Space-sensitive Customization and Semantic Preservation
by: Wang, Qilin, et al.
Published: (2024)
by: Wang, Qilin, et al.
Published: (2024)
MDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture Generation
by: Mao, Xiaofeng, et al.
Published: (2024)
by: Mao, Xiaofeng, et al.
Published: (2024)
MambaGesture: Enhancing Co-Speech Gesture Generation with Mamba and Disentangled Multi-Modality Fusion
by: Fu, Chencan, et al.
Published: (2024)
by: Fu, Chencan, et al.
Published: (2024)
PiT: Progressive Diffusion Transformer
by: Wu, Jiafu, et al.
Published: (2025)
by: Wu, Jiafu, et al.
Published: (2025)
SVP: Style-Enhanced Vivid Portrait Talking Head Diffusion Model
by: Tan, Weipeng, et al.
Published: (2024)
by: Tan, Weipeng, et al.
Published: (2024)
StrandDesigner: Towards Practical Strand Generation with Sketch Guidance
by: Zhang, Na, et al.
Published: (2025)
by: Zhang, Na, et al.
Published: (2025)
OSV: One Step is Enough for High-Quality Image to Video Generation
by: Mao, Xiaofeng, et al.
Published: (2024)
by: Mao, Xiaofeng, et al.
Published: (2024)
Identity-Preserving Text-to-Video Generation Guided by Simple yet Effective Spatial-Temporal Decoupled Representations
by: Wang, Yuji, et al.
Published: (2025)
by: Wang, Yuji, et al.
Published: (2025)
Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10$\times$
by: Zhang, Jiangning, et al.
Published: (2025)
by: Zhang, Jiangning, et al.
Published: (2025)
VividAnimator: An End-to-End Audio and Pose-driven Half-Body Human Animation Framework
by: Huang, Donglin, et al.
Published: (2025)
by: Huang, Donglin, et al.
Published: (2025)
AnomalyDiffusion: Few-Shot Anomaly Image Generation with Diffusion Model
by: Hu, Teng, et al.
Published: (2023)
by: Hu, Teng, et al.
Published: (2023)
Semantic Frame Interpolation
by: Hong, Yijia, et al.
Published: (2025)
by: Hong, Yijia, et al.
Published: (2025)
FFP-300K: Scaling First-Frame Propagation for Generalizable Video Editing
by: Huang, Xijie, et al.
Published: (2026)
by: Huang, Xijie, et al.
Published: (2026)
HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
by: Cai, Yuxuan, et al.
Published: (2025)
by: Cai, Yuxuan, et al.
Published: (2025)
FreeMotion: A Unified Framework for Number-free Text-to-Motion Synthesis
by: Fan, Ke, et al.
Published: (2024)
by: Fan, Ke, et al.
Published: (2024)
FitDiT: Advancing the Authentic Garment Details for High-fidelity Virtual Try-on
by: Jiang, Boyuan, et al.
Published: (2024)
by: Jiang, Boyuan, et al.
Published: (2024)
UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
by: Wang, Xiang, et al.
Published: (2025)
by: Wang, Xiang, et al.
Published: (2025)
PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation
by: Wang, Xiang, et al.
Published: (2024)
by: Wang, Xiang, et al.
Published: (2024)
UniM-OV3D: Uni-Modality Open-Vocabulary 3D Scene Understanding with Fine-Grained Feature Representation
by: He, Qingdong, et al.
Published: (2024)
by: He, Qingdong, et al.
Published: (2024)
TexDreamer: Towards Zero-Shot High-Fidelity 3D Human Texture Generation
by: Liu, Yufei, et al.
Published: (2024)
by: Liu, Yufei, et al.
Published: (2024)
PVG: Progressive Vision Graph for Vision Recognition
by: Wu, Jiafu, et al.
Published: (2023)
by: Wu, Jiafu, et al.
Published: (2023)
Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer
by: Cui, Jiahao, et al.
Published: (2024)
by: Cui, Jiahao, et al.
Published: (2024)
VividFace: Real-Time and Realistic Facial Expression Shadowing for Humanoid Robots
by: Li, Peizhen, et al.
Published: (2026)
by: Li, Peizhen, et al.
Published: (2026)
One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer
by: Shi, Shijun, et al.
Published: (2025)
by: Shi, Shijun, et al.
Published: (2025)
Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
by: Zhang, Jiangning, et al.
Published: (2025)
by: Zhang, Jiangning, et al.
Published: (2025)
DisPose: Disentangling Pose Guidance for Controllable Human Image Animation
by: Li, Hongxiang, et al.
Published: (2024)
by: Li, Hongxiang, et al.
Published: (2024)
AdapNet: Adaptive Noise-Based Network for Low-Quality Image Retrieval
by: Zhang, Sihe, et al.
Published: (2024)
by: Zhang, Sihe, et al.
Published: (2024)
EATFormer: Improving Vision Transformer Inspired by Evolutionary Algorithm
by: Zhang, Jiangning, et al.
Published: (2022)
by: Zhang, Jiangning, et al.
Published: (2022)
PointRWKV: Efficient RWKV-Like Model for Hierarchical Point Cloud Learning
by: He, Qingdong, et al.
Published: (2024)
by: He, Qingdong, et al.
Published: (2024)
What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing
by: Lin, Hangyu, et al.
Published: (2026)
by: Lin, Hangyu, et al.
Published: (2026)
StableAnimator++: Overcoming Pose Misalignment and Face Distortion for Human Image Animation
by: Tu, Shuyuan, et al.
Published: (2025)
by: Tu, Shuyuan, et al.
Published: (2025)
InterAnimate: Taming Region-aware Diffusion Model for Realistic Human Interaction Animation
by: Lin, Yukang, et al.
Published: (2025)
by: Lin, Yukang, et al.
Published: (2025)
When Preferences Diverge: Aligning Diffusion Models with Minority-Aware Adaptive DPO
by: Zhang, Lingfan, et al.
Published: (2025)
by: Zhang, Lingfan, et al.
Published: (2025)
DMAD: Dual Memory Bank for Real-World Anomaly Detection
by: Hu, Jianlong, et al.
Published: (2024)
by: Hu, Jianlong, et al.
Published: (2024)
ID-Sculpt: ID-aware 3D Head Generation from Single In-the-wild Portrait Image
by: Hao, Jinkun, et al.
Published: (2024)
by: Hao, Jinkun, et al.
Published: (2024)
CLIP-AD: A Language-Guided Staged Dual-Path Model for Zero-shot Anomaly Detection
by: Chen, Xuhai, et al.
Published: (2023)
by: Chen, Xuhai, et al.
Published: (2023)
Similar Items
-
SwiftVideo: A Unified Framework for Few-Step Video Generation through Trajectory-Distribution Alignment
by: Sun, Yanxiao, et al.
Published: (2025) -
Collaborative Face Experts Fusion in Video Generation: Boosting Identity Consistency Across Large Face Poses
by: Wang, Yuji, et al.
Published: (2025) -
ArtWeaver: Advanced Dynamic Style Integration via Diffusion Model
by: Xu, Chengming, et al.
Published: (2024) -
DiffFAE: Advancing High-fidelity One-shot Facial Appearance Editing with Space-sensitive Customization and Semantic Preservation
by: Wang, Qilin, et al.
Published: (2024) -
MDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture Generation
by: Mao, Xiaofeng, et al.
Published: (2024)