InterDyn: Controllable Interactive Dynamics with Video Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Akkerman, Rick, Feng, Haiwen, Black, Michael J., Tzionas, Dimitrios, Abrevaya, Victoria Fernández
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916673775206400
author Akkerman, Rick
Feng, Haiwen
Black, Michael J.
Tzionas, Dimitrios
Abrevaya, Victoria Fernández
author_facet Akkerman, Rick
Feng, Haiwen
Black, Michael J.
Tzionas, Dimitrios
Abrevaya, Victoria Fernández
contents Predicting the dynamics of interacting objects is essential for both humans and intelligent systems. However, existing approaches are limited to simplified, toy settings and lack generalizability to complex, real-world environments. Recent advances in generative models have enabled the prediction of state transitions based on interventions, but focus on generating a single future state which neglects the continuous dynamics resulting from the interaction. To address this gap, we propose InterDyn, a novel framework that generates videos of interactive dynamics given an initial frame and a control signal encoding the motion of a driving object or actor. Our key insight is that large video generation models can act as both neural renderers and implicit physics ``simulators'', having learned interactive dynamics from large-scale video data. To effectively harness this capability, we introduce an interactive control mechanism that conditions the video generation process on the motion of the driving entity. Qualitative results demonstrate that InterDyn generates plausible, temporally consistent videos of complex object interactions while generalizing to unseen objects. Quantitative evaluations show that InterDyn outperforms baselines that focus on static state transitions. This work highlights the potential of leveraging video generative models as implicit physics engines. Project page: https://interdyn.is.tue.mpg.de/
format Preprint
id arxiv_https___arxiv_org_abs_2412_11785
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle InterDyn: Controllable Interactive Dynamics with Video Diffusion Models
Akkerman, Rick
Feng, Haiwen
Black, Michael J.
Tzionas, Dimitrios
Abrevaya, Victoria Fernández
Computer Vision and Pattern Recognition
Predicting the dynamics of interacting objects is essential for both humans and intelligent systems. However, existing approaches are limited to simplified, toy settings and lack generalizability to complex, real-world environments. Recent advances in generative models have enabled the prediction of state transitions based on interventions, but focus on generating a single future state which neglects the continuous dynamics resulting from the interaction. To address this gap, we propose InterDyn, a novel framework that generates videos of interactive dynamics given an initial frame and a control signal encoding the motion of a driving object or actor. Our key insight is that large video generation models can act as both neural renderers and implicit physics ``simulators'', having learned interactive dynamics from large-scale video data. To effectively harness this capability, we introduce an interactive control mechanism that conditions the video generation process on the motion of the driving entity. Qualitative results demonstrate that InterDyn generates plausible, temporally consistent videos of complex object interactions while generalizing to unseen objects. Quantitative evaluations show that InterDyn outperforms baselines that focus on static state transitions. This work highlights the potential of leveraging video generative models as implicit physics engines. Project page: https://interdyn.is.tue.mpg.de/
title InterDyn: Controllable Interactive Dynamics with Video Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.11785