MotionGrounder: Grounded Multi-Object Motion Transfer via Diffusion Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Teodoro, Samuel, Chen, Yun, Gunawan, Agus, Kim, Soo Ye, Oh, Jihyong, Kim, Munchurl |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PRIMEdit: Probability Redistribution for Instance-aware Multi-object Video Editing with Benchmark Dataset
by: Teodoro, Samuel, et al.
Published: (2024)
by: Teodoro, Samuel, et al.
Published: (2024)
OmniText: A Training-Free Generalist for Controllable Text-Image Manipulation
by: Gunawan, Agus, et al.
Published: (2025)
by: Gunawan, Agus, et al.
Published: (2025)
BiM-VFI: Bidirectional Motion Field-Guided Frame Interpolation for Video with Non-uniform Motions
by: Seo, Wonyong, et al.
Published: (2024)
by: Seo, Wonyong, et al.
Published: (2024)
FMA-Net++: Motion- and Exposure-Aware Real-World Joint Video Super-Resolution and Deblurring
by: Youk, Geunhyuk, et al.
Published: (2025)
by: Youk, Geunhyuk, et al.
Published: (2025)
FMA-Net: Flow-Guided Dynamic Filtering and Iterative Feature Refinement with Multi-Attention for Joint Video Super-Resolution and Deblurring
by: Youk, Geunhyuk, et al.
Published: (2024)
by: Youk, Geunhyuk, et al.
Published: (2024)
MoBluRF: Motion Deblurring Neural Radiance Fields for Blurry Monocular Video
by: Bui, Minh-Quan Viet, et al.
Published: (2023)
by: Bui, Minh-Quan Viet, et al.
Published: (2023)
Chameleon: Style-Content Disentangled Framework for Cross-Domain Object Compositing
by: Ko, Sukhun, et al.
Published: (2026)
by: Ko, Sukhun, et al.
Published: (2026)
MoDec-GS: Global-to-Local Motion Decomposition and Temporal Interval Adjustment for Compact Dynamic 3D Gaussian Splatting
by: Kwak, Sangwoon, et al.
Published: (2025)
by: Kwak, Sangwoon, et al.
Published: (2025)
MoBGS: Motion Deblurring Dynamic 3D Gaussian Splatting for Blurry Monocular Video
by: Bui, Minh-Quan Viet, et al.
Published: (2025)
by: Bui, Minh-Quan Viet, et al.
Published: (2025)
SplineGS: Robust Motion-Adaptive Spline for Real-Time Dynamic 3D Gaussians from Monocular Video
by: Park, Jongmin, et al.
Published: (2024)
by: Park, Jongmin, et al.
Published: (2024)
PropFly: Learning to Propagate via On-the-Fly Supervision from Pre-trained Video Diffusion Models
by: Seo, Wonyong, et al.
Published: (2026)
by: Seo, Wonyong, et al.
Published: (2026)
MultiMotion: Multi Subject Video Motion Transfer via Video Diffusion Transformer
by: Liu, Penghui, et al.
Published: (2025)
by: Liu, Penghui, et al.
Published: (2025)
EcoSplat: Efficiency-controllable Feed-forward 3D Gaussian Splatting from Multi-view Images
by: Park, Jongmin, et al.
Published: (2025)
by: Park, Jongmin, et al.
Published: (2025)
SkateFormer: Skeletal-Temporal Transformer for Human Action Recognition
by: Do, Jeonghyeok, et al.
Published: (2024)
by: Do, Jeonghyeok, et al.
Published: (2024)
MoRel: Long-Range Flicker-Free 4D Motion Modeling via Anchor Relay-based Bidirectional Blending with Hierarchical Densification
by: Kwak, Sangwoon, et al.
Published: (2025)
by: Kwak, Sangwoon, et al.
Published: (2025)
Bridging the Skeleton-Text Modality Gap: Diffusion-Powered Modality Alignment for Zero-shot Skeleton-based Action Recognition
by: Do, Jeonghyeok, et al.
Published: (2024)
by: Do, Jeonghyeok, et al.
Published: (2024)
Let Your Image Move with Your Motion! -- Implicit Multi-Object Multi-Motion Transfer
by: Li, Yuze, et al.
Published: (2026)
by: Li, Yuze, et al.
Published: (2026)
Leveraging Prior Knowledge of Diffusion Model for Person Search
by: Kim, Giyeol, et al.
Published: (2025)
by: Kim, Giyeol, et al.
Published: (2025)
C-DiffSET: Leveraging Latent Diffusion for SAR-to-EO Image Translation with Confidence-Guided Reliable Object Generation
by: Do, Jeonghyeok, et al.
Published: (2024)
by: Do, Jeonghyeok, et al.
Published: (2024)
Motion-Aware Transformer for Multi-Object Tracking
by: Yang, Xu, et al.
Published: (2025)
by: Yang, Xu, et al.
Published: (2025)
Less is More: Decoder-Free Masked Modeling for Efficient Skeleton Representation Learning
by: Do, Jeonghyeok, et al.
Published: (2026)
by: Do, Jeonghyeok, et al.
Published: (2026)
FRAMER: Frequency-Aligned Self-Distillation with Adaptive Modulation Leveraging Diffusion Priors for Real-World Image Super-Resolution
by: Choi, Seungho, et al.
Published: (2025)
by: Choi, Seungho, et al.
Published: (2025)
One Look is Enough: Seamless Patchwise Refinement for Zero-Shot Monocular Depth Estimation on High-Resolution Images
by: Kwon, Byeongjun, et al.
Published: (2025)
by: Kwon, Byeongjun, et al.
Published: (2025)
DiTTo: Scalable Order-aware All-in-One Image Restoration Agent
by: Choi, Seungho, et al.
Published: (2026)
by: Choi, Seungho, et al.
Published: (2026)
EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer
by: Yang, Yuxiao, et al.
Published: (2025)
by: Yang, Yuxiao, et al.
Published: (2025)
Diffusion-based Data Augmentation and Knowledge Distillation with Generated Soft Labels Solving Data Scarcity Problems of SAR Oil Spill Segmentation
by: Moon, Jaeho, et al.
Published: (2024)
by: Moon, Jaeho, et al.
Published: (2024)
Video Motion Transfer with Diffusion Transformers
by: Pondaven, Alexander, et al.
Published: (2024)
by: Pondaven, Alexander, et al.
Published: (2024)
MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video Generation
by: Liu, Yanchen, et al.
Published: (2025)
by: Liu, Yanchen, et al.
Published: (2025)
PlugTrack: Multi-Perceptive Motion Analysis for Adaptive Fusion in Multi-Object Tracking
by: Kim, Seungjae, et al.
Published: (2025)
by: Kim, Seungjae, et al.
Published: (2025)
Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion Transfer
by: Shi, Qingyu, et al.
Published: (2025)
by: Shi, Qingyu, et al.
Published: (2025)
MoCHA-former: Moiré-Conditioned Hybrid Adaptive Transformer for Video Demoiréing
by: Sung, Jeahun, et al.
Published: (2025)
by: Sung, Jeahun, et al.
Published: (2025)
Tri-Prompting: Video Diffusion with Unified Control over Scene, Subject, and Motion
by: Zhou, Zhenghong, et al.
Published: (2026)
by: Zhou, Zhenghong, et al.
Published: (2026)
AM-SORT: Adaptable Motion Predictor with Historical Trajectory Embedding for Multi-Object Tracking
by: Kim, Vitaliy, et al.
Published: (2024)
by: Kim, Vitaliy, et al.
Published: (2024)
IF-MDM: Implicit Face Motion Diffusion Model for High-Fidelity Realtime Talking Head Generation
by: Yang, Sejong, et al.
Published: (2024)
by: Yang, Sejong, et al.
Published: (2024)
Dynamic Full-body Motion Agent with Object Interaction via Blending Pre-trained Modular Controllers
by: Nam, Sanghyeok, et al.
Published: (2026)
by: Nam, Sanghyeok, et al.
Published: (2026)
Motion2Motion: Cross-topology Motion Transfer with Sparse Correspondence
by: Chen, Ling-Hao, et al.
Published: (2025)
by: Chen, Ling-Hao, et al.
Published: (2025)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
MotionCFG: Boosting Motion Dynamics via Stochastic Concept Perturbation
by: Kim, Byungjun, et al.
Published: (2026)
by: Kim, Byungjun, et al.
Published: (2026)
Spectral Motion Alignment for Video Motion Transfer using Diffusion Models
by: Park, Geon Yeong, et al.
Published: (2024)
by: Park, Geon Yeong, et al.
Published: (2024)
MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation
by: Shi, Shuwei, et al.
Published: (2024)
by: Shi, Shuwei, et al.
Published: (2024)
Similar Items
-
PRIMEdit: Probability Redistribution for Instance-aware Multi-object Video Editing with Benchmark Dataset
by: Teodoro, Samuel, et al.
Published: (2024) -
OmniText: A Training-Free Generalist for Controllable Text-Image Manipulation
by: Gunawan, Agus, et al.
Published: (2025) -
BiM-VFI: Bidirectional Motion Field-Guided Frame Interpolation for Video with Non-uniform Motions
by: Seo, Wonyong, et al.
Published: (2024) -
FMA-Net++: Motion- and Exposure-Aware Real-World Joint Video Super-Resolution and Deblurring
by: Youk, Geunhyuk, et al.
Published: (2025) -
FMA-Net: Flow-Guided Dynamic Filtering and Iterative Feature Refinement with Multi-Attention for Joint Video Super-Resolution and Deblurring
by: Youk, Geunhyuk, et al.
Published: (2024)