Consistent Human Image and Video Generation with Spatially Conditioned Diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Mingdeng, Mou, Chong, Yuan, Ziyang, Wang, Xintao, Zhang, Zhaoyang, Shan, Ying, Zheng, Yinqiang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ReVideo: Remake a Video with Motion and Content Control
by: Mou, Chong, et al.
Published: (2024)
by: Mou, Chong, et al.
Published: (2024)
MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model
by: Niu, Muyao, et al.
Published: (2024)
by: Niu, Muyao, et al.
Published: (2024)
Instruction-based Image Manipulation by Watching How Things Move
by: Cao, Mingdeng, et al.
Published: (2024)
by: Cao, Mingdeng, et al.
Published: (2024)
DiffEditor: Boosting Accuracy and Flexibility on Diffusion-based Image Editing
by: Mou, Chong, et al.
Published: (2024)
by: Mou, Chong, et al.
Published: (2024)
AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models
by: Niu, Muyao, et al.
Published: (2025)
by: Niu, Muyao, et al.
Published: (2025)
Rolling Shutter Correction with Intermediate Distortion Flow Estimation
by: Cao, Mingdeng, et al.
Published: (2024)
by: Cao, Mingdeng, et al.
Published: (2024)
Image Conductor: Precision Control for Interactive Video Synthesis
by: Li, Yaowei, et al.
Published: (2024)
by: Li, Yaowei, et al.
Published: (2024)
VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control
by: Bian, Yuxuan, et al.
Published: (2025)
by: Bian, Yuxuan, et al.
Published: (2025)
MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions
by: Ju, Xuan, et al.
Published: (2024)
by: Ju, Xuan, et al.
Published: (2024)
360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion Model
by: Wang, Qian, et al.
Published: (2024)
by: Wang, Qian, et al.
Published: (2024)
LaMD: Latent Motion Diffusion for Image-Conditional Video Generation
by: Hu, Yaosi, et al.
Published: (2023)
by: Hu, Yaosi, et al.
Published: (2023)
All-in-One Transferring Image Compression from Human Perception to Multi-Machine Perception
by: Zhao, Jiancheng, et al.
Published: (2025)
by: Zhao, Jiancheng, et al.
Published: (2025)
VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
by: Chen, Haoxin, et al.
Published: (2024)
by: Chen, Haoxin, et al.
Published: (2024)
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
by: Wang, Zhouxia, et al.
Published: (2023)
by: Wang, Zhouxia, et al.
Published: (2023)
Analysis and Benchmarking of Extending Blind Face Image Restoration to Videos
by: Wang, Zhouxia, et al.
Published: (2024)
by: Wang, Zhouxia, et al.
Published: (2024)
EventHDR: from Event to High-Speed HDR Videos and Beyond
by: Zou, Yunhao, et al.
Published: (2024)
by: Zou, Yunhao, et al.
Published: (2024)
Hierarchical Codec Diffusion for Video-to-Speech Generation
by: Ye, Jiaxin, et al.
Published: (2026)
by: Ye, Jiaxin, et al.
Published: (2026)
FreeNoise: Tuning-Free Longer Video Diffusion via Noise Rescheduling
by: Qiu, Haonan, et al.
Published: (2023)
by: Qiu, Haonan, et al.
Published: (2023)
MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation
by: Shi, Shuwei, et al.
Published: (2024)
by: Shi, Shuwei, et al.
Published: (2024)
CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation
by: Wang, Xiang, et al.
Published: (2024)
by: Wang, Xiang, et al.
Published: (2024)
ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning
by: Huang, Yuzhou, et al.
Published: (2025)
by: Huang, Yuzhou, et al.
Published: (2025)
ContextAnyone: Context-Aware Diffusion for Character-Consistent Text-to-Video Generation
by: Mai, Ziyang, et al.
Published: (2025)
by: Mai, Ziyang, et al.
Published: (2025)
BrushNet: A Plug-and-Play Image Inpainting Model with Decomposed Dual-Branch Diffusion
by: Ju, Xuan, et al.
Published: (2024)
by: Ju, Xuan, et al.
Published: (2024)
DepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth Estimation
by: Dong, Yue-Jiang, et al.
Published: (2025)
by: Dong, Yue-Jiang, et al.
Published: (2025)
StyleAdapter: A Unified Stylized Image Generation Model
by: Wang, Zhouxia, et al.
Published: (2023)
by: Wang, Zhouxia, et al.
Published: (2023)
OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation
by: Li, Weiqi, et al.
Published: (2024)
by: Li, Weiqi, et al.
Published: (2024)
Diffusion-Based Hierarchical Image Steganography
by: Xu, Youmin, et al.
Published: (2024)
by: Xu, Youmin, et al.
Published: (2024)
Measuring 3D Spatial Geometric Consistency in Dynamic Generated Videos
by: Dou, Weijia, et al.
Published: (2026)
by: Dou, Weijia, et al.
Published: (2026)
FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers
by: He, Xuanhua, et al.
Published: (2025)
by: He, Xuanhua, et al.
Published: (2025)
Edit Temporal-Consistent Videos with Image Diffusion Model
by: Wang, Yuanzhi, et al.
Published: (2023)
by: Wang, Yuanzhi, et al.
Published: (2023)
VividPose: Advancing Stable Video Diffusion for Realistic Human Image Animation
by: Wang, Qilin, et al.
Published: (2024)
by: Wang, Qilin, et al.
Published: (2024)
Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
by: Yu, Jiwen, et al.
Published: (2025)
by: Yu, Jiwen, et al.
Published: (2025)
Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos
by: Ma, Yue, et al.
Published: (2023)
by: Ma, Yue, et al.
Published: (2023)
Enhancing Shape Perception and Segmentation Consistency for Industrial Image Inspection
by: Mao, Guoxuan, et al.
Published: (2025)
by: Mao, Guoxuan, et al.
Published: (2025)
SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints
by: Bai, Jianhong, et al.
Published: (2024)
by: Bai, Jianhong, et al.
Published: (2024)
Towards Spatially Consistent Image Generation: On Incorporating Intrinsic Scene Properties into Diffusion Models
by: Lee, Hyundo, et al.
Published: (2025)
by: Lee, Hyundo, et al.
Published: (2025)
Diffusion4D: Fast Spatial-temporal Consistent 4D Generation via Video Diffusion Models
by: Liang, Hanwen, et al.
Published: (2024)
by: Liang, Hanwen, et al.
Published: (2024)
InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models
by: Xu, Jiale, et al.
Published: (2024)
by: Xu, Jiale, et al.
Published: (2024)
JVID: Joint Video-Image Diffusion for Visual-Quality and Temporal-Consistency in Video Generation
by: Reynaud, Hadrien, et al.
Published: (2024)
by: Reynaud, Hadrien, et al.
Published: (2024)
Similar Items
-
ReVideo: Remake a Video with Motion and Content Control
by: Mou, Chong, et al.
Published: (2024) -
MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model
by: Niu, Muyao, et al.
Published: (2024) -
Instruction-based Image Manipulation by Watching How Things Move
by: Cao, Mingdeng, et al.
Published: (2024) -
DiffEditor: Boosting Accuracy and Flexibility on Diffusion-based Image Editing
by: Mou, Chong, et al.
Published: (2024) -
AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models
by: Niu, Muyao, et al.
Published: (2025)