Disentangling Foreground and Background Motion for Enhanced Realism in Human Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Jinlin, Yu, Kai, Feng, Mengyang, Guo, Xiefan, Cui, Miaomiao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
I4VGen: Image as Free Stepping Stone for Text-to-Video Generation
by: Guo, Xiefan, et al.
Published: (2024)
by: Guo, Xiefan, et al.
Published: (2024)
FB-CLIP: Fine-Grained Zero-Shot Anomaly Detection with Foreground-Background Disentanglement
by: Hu, Ming, et al.
Published: (2026)
by: Hu, Ming, et al.
Published: (2026)
Demystifying Foreground-Background Memorization in Diffusion Models
by: Di, Jimmy Z., et al.
Published: (2025)
by: Di, Jimmy Z., et al.
Published: (2025)
TokenMotion: Decoupled Motion Control via Token Disentanglement for Human-centric Video Generation
by: Li, Ruineng, et al.
Published: (2025)
by: Li, Ruineng, et al.
Published: (2025)
DisentTalk: Cross-lingual Talking Face Generation via Semantic Disentangled Diffusion Model
by: Liu, Kangwei, et al.
Published: (2025)
by: Liu, Kangwei, et al.
Published: (2025)
InitNO: Boosting Text-to-Image Diffusion Models via Initial Noise Optimization
by: Guo, Xiefan, et al.
Published: (2024)
by: Guo, Xiefan, et al.
Published: (2024)
PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
by: Zhan, Yu-Wei, et al.
Published: (2025)
by: Zhan, Yu-Wei, et al.
Published: (2025)
Enhanced Controllability of Diffusion Models via Feature Disentanglement and Realism-Enhanced Sampling Methods
by: Cho, Wonwoong, et al.
Published: (2023)
by: Cho, Wonwoong, et al.
Published: (2023)
Is Visual Realism Enough? Evaluating Gait Biometric Fidelity in Generative AI Human Animation
by: DeAndres-Tame, Ivan, et al.
Published: (2025)
by: DeAndres-Tame, Ivan, et al.
Published: (2025)
BodyMetric: Evaluating the Realism of Human Bodies in Text-to-Image Generation
by: Andreou, Nefeli, et al.
Published: (2024)
by: Andreou, Nefeli, et al.
Published: (2024)
"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
by: Gu, Jing, et al.
Published: (2025)
by: Gu, Jing, et al.
Published: (2025)
ShortFT: Diffusion Model Alignment via Shortcut-based Fine-Tuning
by: Guo, Xiefan, et al.
Published: (2025)
by: Guo, Xiefan, et al.
Published: (2025)
Disentangled Hierarchical VAE for 3D Human-Human Interaction Generation
by: Geng, Zichen, et al.
Published: (2026)
by: Geng, Zichen, et al.
Published: (2026)
MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement
by: Deng, Yufan, et al.
Published: (2025)
by: Deng, Yufan, et al.
Published: (2025)
PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
by: Huang, Yidong, et al.
Published: (2026)
by: Huang, Yidong, et al.
Published: (2026)
MUSTAN: Multi-scale Temporal Context as Attention for Robust Video Foreground Segmentation
by: Pokala, Praveen Kumar, et al.
Published: (2024)
by: Pokala, Praveen Kumar, et al.
Published: (2024)
MoViD: View-Invariant 3D Human Pose Estimation via Motion-View Disentanglement
by: Liu, Yejia, et al.
Published: (2026)
by: Liu, Yejia, et al.
Published: (2026)
Pareto-Enhanced Portrait Generation: Vision-Aligned Text Supervision for Alignment, Realism, and Aesthetics
by: Wang, Yunlong, et al.
Published: (2026)
by: Wang, Yunlong, et al.
Published: (2026)
DisCo: Disentangled Control for Realistic Human Dance Generation
by: Wang, Tan, et al.
Published: (2023)
by: Wang, Tan, et al.
Published: (2023)
DiffusionGAN3D: Boosting Text-guided 3D Generation and Domain Adaptation by Combining 3D GANs and Diffusion Priors
by: Lei, Biwen, et al.
Published: (2023)
by: Lei, Biwen, et al.
Published: (2023)
Improving Out-of-Distribution Detection with Disentangled Foreground and Background Features
by: Ding, Choubo, et al.
Published: (2023)
by: Ding, Choubo, et al.
Published: (2023)
Zero-shot Synthetic Video Realism Enhancement via Structure-aware Denoising
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
Fore-Mamba3D: Mamba-based Foreground-Enhanced Encoding for 3D Object Detection
by: Ning, Zhiwei, et al.
Published: (2026)
by: Ning, Zhiwei, et al.
Published: (2026)
Resource-Efficient Motion Control for Video Generation via Dynamic Mask Guidance
by: Feng, Sicong, et al.
Published: (2025)
by: Feng, Sicong, et al.
Published: (2025)
Motion-Aware Caching for Efficient Autoregressive Video Generation
by: Xu, Jing, et al.
Published: (2026)
by: Xu, Jing, et al.
Published: (2026)
Generative AI-Driven High-Fidelity Human Motion Simulation
by: Iyer, Hari, et al.
Published: (2025)
by: Iyer, Hari, et al.
Published: (2025)
Towards Fine-Grained Human Motion Video Captioning
by: Song, Guorui, et al.
Published: (2025)
by: Song, Guorui, et al.
Published: (2025)
Motion Dreamer: Boundary Conditional Motion Reasoning for Physically Coherent Video Generation
by: Xu, Tianshuo, et al.
Published: (2024)
by: Xu, Tianshuo, et al.
Published: (2024)
MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
by: Zhang, Yuang, et al.
Published: (2024)
by: Zhang, Yuang, et al.
Published: (2024)
StickMotion: Generating 3D Human Motions by Drawing a Stickman
by: Wang, Tao, et al.
Published: (2025)
by: Wang, Tao, et al.
Published: (2025)
Controllable Video Generation with Provable Disentanglement
by: Shen, Yifan, et al.
Published: (2025)
by: Shen, Yifan, et al.
Published: (2025)
Motion Guided Token Compression for Efficient Masked Video Modeling
by: Feng, Yukun, et al.
Published: (2024)
by: Feng, Yukun, et al.
Published: (2024)
DEMO: Disentangled Motion Latent Flow Matching for Fine-Grained Controllable Talking Portrait Synthesis
by: Chen, Peiyin, et al.
Published: (2025)
by: Chen, Peiyin, et al.
Published: (2025)
Disentangling Foreground and Background for vision-Language Navigation via Online Augmentation
by: Xu, Yunbo, et al.
Published: (2025)
by: Xu, Yunbo, et al.
Published: (2025)
RAGME: Retrieval Augmented Video Generation for Enhanced Motion Realism
by: Peruzzo, Elia, et al.
Published: (2025)
by: Peruzzo, Elia, et al.
Published: (2025)
CMD-HAR: Cross-Modal Disentanglement for Wearable Human Activity Recognition
by: Yu, Ying, et al.
Published: (2025)
by: Yu, Ying, et al.
Published: (2025)
FreeOrbit4D: Training-Free Arbitrary Camera Redirection for Monocular Videos via Foreground-Complete 4D Reconstruction
by: Cao, Wei, et al.
Published: (2026)
by: Cao, Wei, et al.
Published: (2026)
An Evaluation Framework for Product Images Background Inpainting based on Human Feedback and Product Consistency
by: Liang, Yuqi, et al.
Published: (2024)
by: Liang, Yuqi, et al.
Published: (2024)
WANDR: Intention-guided Human Motion Generation
by: Diomataris, Markos, et al.
Published: (2024)
by: Diomataris, Markos, et al.
Published: (2024)
STANCE: Motion Coherent Video Generation Via Sparse-to-Dense Anchored Encoding
by: Chen, Zhifei, et al.
Published: (2025)
by: Chen, Zhifei, et al.
Published: (2025)
Similar Items
-
I4VGen: Image as Free Stepping Stone for Text-to-Video Generation
by: Guo, Xiefan, et al.
Published: (2024) -
FB-CLIP: Fine-Grained Zero-Shot Anomaly Detection with Foreground-Background Disentanglement
by: Hu, Ming, et al.
Published: (2026) -
Demystifying Foreground-Background Memorization in Diffusion Models
by: Di, Jimmy Z., et al.
Published: (2025) -
TokenMotion: Decoupled Motion Control via Token Disentanglement for Human-centric Video Generation
by: Li, Ruineng, et al.
Published: (2025) -
DisentTalk: Cross-lingual Talking Face Generation via Semantic Disentangled Diffusion Model
by: Liu, Kangwei, et al.
Published: (2025)