Learning to Generate Rigid Body Interactions with Video Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Romero, David, Bermudez, Ariana, Iablochnikov, Viacheslav, Li, Hao, Pizzati, Fabio, Laptev, Ivan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference
by: Yuan, Jianhao, et al.
Published: (2025)
by: Yuan, Jianhao, et al.
Published: (2025)
ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation
by: Khalifi, Omar El, et al.
Published: (2026)
by: Khalifi, Omar El, et al.
Published: (2026)
MessyKitchens: Contact-rich object-level 3D scene reconstruction
by: Ansari, Junaid Ahmed, et al.
Published: (2026)
by: Ansari, Junaid Ahmed, et al.
Published: (2026)
Video Motion Transfer with Diffusion Transformers
by: Pondaven, Alexander, et al.
Published: (2024)
by: Pondaven, Alexander, et al.
Published: (2024)
PhysMoDPO: Physically-Plausible Humanoid Motion with Preference Optimization
by: Zhang, Yangsong, et al.
Published: (2026)
by: Zhang, Yangsong, et al.
Published: (2026)
PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation
by: Liu, Shaowei, et al.
Published: (2024)
by: Liu, Shaowei, et al.
Published: (2024)
ActionParty: Multi-Subject Action Binding in Generative Video Games
by: Pondaven, Alexander, et al.
Published: (2026)
by: Pondaven, Alexander, et al.
Published: (2026)
MatchDiffusion: Training-free Generation of Match-cuts
by: Pardo, Alejandro, et al.
Published: (2024)
by: Pardo, Alejandro, et al.
Published: (2024)
Attacks on multimodal models
by: Iablochnikov, Viacheslav, et al.
Published: (2024)
by: Iablochnikov, Viacheslav, et al.
Published: (2024)
Interactive Generation of Laparoscopic Videos with Diffusion Models
by: Iliash, Ivan, et al.
Published: (2024)
by: Iliash, Ivan, et al.
Published: (2024)
Specify and Edit: Overcoming Ambiguity in Text-Based Image Editing
by: Iakovleva, Ekaterina, et al.
Published: (2024)
by: Iakovleva, Ekaterina, et al.
Published: (2024)
Latent Guard: a Safety Framework for Text-to-image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
Visual Language Models as Zero-Shot Deepfake Detectors
by: Pirogov, Viacheslav
Published: (2025)
by: Pirogov, Viacheslav
Published: (2025)
On Pretraining Data Diversity for Self-Supervised Learning
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation
by: Jeong, Hyeonho, et al.
Published: (2024)
by: Jeong, Hyeonho, et al.
Published: (2024)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models
by: Surkov, Viacheslav, et al.
Published: (2024)
by: Surkov, Viacheslav, et al.
Published: (2024)
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
Contextualized Diffusion Models for Text-Guided Image and Video Generation
by: Yang, Ling, et al.
Published: (2024)
by: Yang, Ling, et al.
Published: (2024)
VideoPDE: Unified Generative PDE Solving via Video Inpainting Diffusion Models
by: Li, Edward, et al.
Published: (2025)
by: Li, Edward, et al.
Published: (2025)
MirrorCheck: Efficient Adversarial Defense for Vision-Language Models
by: Fares, Samar, et al.
Published: (2024)
by: Fares, Samar, et al.
Published: (2024)
Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation
by: Lin, Shanchuan, et al.
Published: (2025)
by: Lin, Shanchuan, et al.
Published: (2025)
VideoDPO: Omni-Preference Alignment for Video Diffusion Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation
by: Arkhipkin, Vladimir, et al.
Published: (2025)
by: Arkhipkin, Vladimir, et al.
Published: (2025)
Long Story Short: Story-level Video Understanding from 20K Short Films
by: Ghermi, Ridouane, et al.
Published: (2024)
by: Ghermi, Ridouane, et al.
Published: (2024)
Continual Learning of Diffusion Models with Generative Distillation
by: Masip, Sergi, et al.
Published: (2023)
by: Masip, Sergi, et al.
Published: (2023)
Evaluating Deepfake Detectors in the Wild
by: Pirogov, Viacheslav, et al.
Published: (2025)
by: Pirogov, Viacheslav, et al.
Published: (2025)
A Survey on Video Diffusion Models
by: Xing, Zhen, et al.
Published: (2023)
by: Xing, Zhen, et al.
Published: (2023)
VRAG: Learning World Models for Interactive Video Generation
by: Chen, Taiye, et al.
Published: (2025)
by: Chen, Taiye, et al.
Published: (2025)
SyncDiff: Synchronized Motion Diffusion for Multi-Body Human-Object Interaction Synthesis
by: He, Wenkun, et al.
Published: (2024)
by: He, Wenkun, et al.
Published: (2024)
AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video Generation
by: Girish, Sharath, et al.
Published: (2025)
by: Girish, Sharath, et al.
Published: (2025)
TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
Warped Diffusion: Solving Video Inverse Problems with Image Diffusion Models
by: Daras, Giannis, et al.
Published: (2024)
by: Daras, Giannis, et al.
Published: (2024)
Reproducing DragDiffusion: Interactive Point-Based Editing with Diffusion Models
by: Subhan, Ali, et al.
Published: (2026)
by: Subhan, Ali, et al.
Published: (2026)
Diffusion Counterfactual Generation with Semantic Abduction
by: Rasal, Rajat, et al.
Published: (2025)
by: Rasal, Rajat, et al.
Published: (2025)
Diffusion Adversarial Post-Training for One-Step Video Generation
by: Lin, Shanchuan, et al.
Published: (2025)
by: Lin, Shanchuan, et al.
Published: (2025)
PoGDiff: Product-of-Gaussians Diffusion Models for Imbalanced Text-to-Image Generation
by: Wang, Ziyan, et al.
Published: (2025)
by: Wang, Ziyan, et al.
Published: (2025)
RigidFormer: Learning Rigid Dynamics using Transformers
by: Dou, Zhiyang, et al.
Published: (2026)
by: Dou, Zhiyang, et al.
Published: (2026)
Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback
by: Furuta, Hiroki, et al.
Published: (2024)
by: Furuta, Hiroki, et al.
Published: (2024)
Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Similar Items
-
LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference
by: Yuan, Jianhao, et al.
Published: (2025) -
ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation
by: Khalifi, Omar El, et al.
Published: (2026) -
MessyKitchens: Contact-rich object-level 3D scene reconstruction
by: Ansari, Junaid Ahmed, et al.
Published: (2026) -
Video Motion Transfer with Diffusion Transformers
by: Pondaven, Alexander, et al.
Published: (2024) -
PhysMoDPO: Physically-Plausible Humanoid Motion with Preference Optimization
by: Zhang, Yangsong, et al.
Published: (2026)