MotionGrounder: Grounded Multi-Object Motion Transfer via Diffusion Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Teodoro, Samuel, Chen, Yun, Gunawan, Agus, Kim, Soo Ye, Oh, Jihyong, Kim, Munchurl
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914437328273408
author Teodoro, Samuel
Chen, Yun
Gunawan, Agus
Kim, Soo Ye
Oh, Jihyong
Kim, Munchurl
author_facet Teodoro, Samuel
Chen, Yun
Gunawan, Agus
Kim, Soo Ye
Oh, Jihyong
Kim, Munchurl
contents Motion transfer enables controllable video generation by transferring temporal dynamics from a reference video to synthesize a new video conditioned on a target caption. However, existing Diffusion Transformer (DiT)-based methods are limited to single-object videos, restricting fine-grained control in real-world scenes with multiple objects. In this work, we introduce MotionGrounder, a DiT-based framework that firstly handles motion transfer with multi-object controllability. Our Flow-based Motion Signal (FMS) in MotionGrounder provides a stable motion prior for target video generation, while our Object-Caption Alignment Loss (OCAL) grounds object captions to their corresponding spatial regions. We further propose a new Object Grounding Score (OGS), which jointly evaluates (i) spatial alignment between source video objects and their generated counterparts and (ii) semantic consistency between each generated object and its target caption. Our experiments show that MotionGrounder consistently outperforms recent baselines across quantitative, qualitative, and human evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2604_00853
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MotionGrounder: Grounded Multi-Object Motion Transfer via Diffusion Transformer
Teodoro, Samuel
Chen, Yun
Gunawan, Agus
Kim, Soo Ye
Oh, Jihyong
Kim, Munchurl
Computer Vision and Pattern Recognition
Motion transfer enables controllable video generation by transferring temporal dynamics from a reference video to synthesize a new video conditioned on a target caption. However, existing Diffusion Transformer (DiT)-based methods are limited to single-object videos, restricting fine-grained control in real-world scenes with multiple objects. In this work, we introduce MotionGrounder, a DiT-based framework that firstly handles motion transfer with multi-object controllability. Our Flow-based Motion Signal (FMS) in MotionGrounder provides a stable motion prior for target video generation, while our Object-Caption Alignment Loss (OCAL) grounds object captions to their corresponding spatial regions. We further propose a new Object Grounding Score (OGS), which jointly evaluates (i) spatial alignment between source video objects and their generated counterparts and (ii) semantic consistency between each generated object and its target caption. Our experiments show that MotionGrounder consistently outperforms recent baselines across quantitative, qualitative, and human evaluations.
title MotionGrounder: Grounded Multi-Object Motion Transfer via Diffusion Transformer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.00853