Learning to Generate Rigid Body Interactions with Video Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Romero, David, Bermudez, Ariana, Iablochnikov, Viacheslav, Li, Hao, Pizzati, Fabio, Laptev, Ivan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908903057391616
author Romero, David
Bermudez, Ariana
Iablochnikov, Viacheslav
Li, Hao
Pizzati, Fabio
Laptev, Ivan
author_facet Romero, David
Bermudez, Ariana
Iablochnikov, Viacheslav
Li, Hao
Pizzati, Fabio
Laptev, Ivan
contents Recent video generation models have achieved remarkable progress and are now deployed in film, social media production, and advertising. Beyond their creative potential, such models also hold promise as world simulators for robotics and embodied decision making. Despite strong advances, current approaches still struggle to generate physically plausible object interactions and lack object-level control mechanisms. To address these limitations, we introduce KineMask, an approach for video generation that enables realistic rigid body control, interactions, and effects. Given a single image and a specified object velocity, our method generates videos with inferred motions and future object interactions. We propose a two-stage training strategy that gradually removes future motion supervision via object masks. Using this strategy we train video diffusion models (VDMs) on synthetic scenes of simple interactions and demonstrate significant improvements and generalization to rigid body and hand-object interactions in real scenes. Furthermore, KineMask integrates low-level motion control with high-level textual conditioning via predicted scene descriptions, leading to support for synthesis of complex dynamical phenomena. Our experiments show that KineMask generalizes to different VDMs and achieves strong improvements over recent models of comparable size. Ablation studies further highlight the complementary roles of low- and high-level conditioning in VDMs. Project Page: https://daromog.github.io/KineMask/
format Preprint
id arxiv_https___arxiv_org_abs_2510_02284
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning to Generate Rigid Body Interactions with Video Diffusion Models
Romero, David
Bermudez, Ariana
Iablochnikov, Viacheslav
Li, Hao
Pizzati, Fabio
Laptev, Ivan
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Recent video generation models have achieved remarkable progress and are now deployed in film, social media production, and advertising. Beyond their creative potential, such models also hold promise as world simulators for robotics and embodied decision making. Despite strong advances, current approaches still struggle to generate physically plausible object interactions and lack object-level control mechanisms. To address these limitations, we introduce KineMask, an approach for video generation that enables realistic rigid body control, interactions, and effects. Given a single image and a specified object velocity, our method generates videos with inferred motions and future object interactions. We propose a two-stage training strategy that gradually removes future motion supervision via object masks. Using this strategy we train video diffusion models (VDMs) on synthetic scenes of simple interactions and demonstrate significant improvements and generalization to rigid body and hand-object interactions in real scenes. Furthermore, KineMask integrates low-level motion control with high-level textual conditioning via predicted scene descriptions, leading to support for synthesis of complex dynamical phenomena. Our experiments show that KineMask generalizes to different VDMs and achieves strong improvements over recent models of comparable size. Ablation studies further highlight the complementary roles of low- and high-level conditioning in VDMs. Project Page: https://daromog.github.io/KineMask/
title Learning to Generate Rigid Body Interactions with Video Diffusion Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.02284