Search2Motion: Training-Free Object-Level Motion Control via Attention-Consensus Search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Sainan, Wu, Tz-Ying, Valdez, Hector A, Tripathi, Subarna
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915872590790656
author Liu, Sainan
Wu, Tz-Ying
Valdez, Hector A
Tripathi, Subarna
author_facet Liu, Sainan
Wu, Tz-Ying
Valdez, Hector A
Tripathi, Subarna
contents We present Search2Motion, a training-free framework for object-level motion editing in image-to-video generation. Unlike prior methods requiring trajectories, bounding boxes, masks, or motion fields, Search2Motion adopts target-frame-based control, leveraging first-last-frame motion priors to realize object relocation while preserving scene stability without fine-tuning. Reliable target-frame construction is achieved through semantic-guided object insertion and robust background inpainting. We further show that early-step self-attention maps predict object and camera dynamics, offering interpretable user feedback and motivating ACE-Seed (Attention Consensus for Early-step Seed selection), a lightweight search strategy that improves motion fidelity without look-ahead sampling or external evaluators. Noting that existing benchmarks conflate object and camera motion, we introduce S2M-DAVIS and S2M-OMB for stable-camera, object-only evaluation, alongside FLF2V-obj metrics that isolate object artifacts without requiring ground-truth trajectories. Search2Motion consistently outperforms baselines on FLF2V-obj and VBench.
format Preprint
id arxiv_https___arxiv_org_abs_2603_16711
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Search2Motion: Training-Free Object-Level Motion Control via Attention-Consensus Search
Liu, Sainan
Wu, Tz-Ying
Valdez, Hector A
Tripathi, Subarna
Computer Vision and Pattern Recognition
We present Search2Motion, a training-free framework for object-level motion editing in image-to-video generation. Unlike prior methods requiring trajectories, bounding boxes, masks, or motion fields, Search2Motion adopts target-frame-based control, leveraging first-last-frame motion priors to realize object relocation while preserving scene stability without fine-tuning. Reliable target-frame construction is achieved through semantic-guided object insertion and robust background inpainting. We further show that early-step self-attention maps predict object and camera dynamics, offering interpretable user feedback and motivating ACE-Seed (Attention Consensus for Early-step Seed selection), a lightweight search strategy that improves motion fidelity without look-ahead sampling or external evaluators. Noting that existing benchmarks conflate object and camera motion, we introduce S2M-DAVIS and S2M-OMB for stable-camera, object-only evaluation, alongside FLF2V-obj metrics that isolate object artifacts without requiring ground-truth trajectories. Search2Motion consistently outperforms baselines on FLF2V-obj and VBench.
title Search2Motion: Training-Free Object-Level Motion Control via Attention-Consensus Search
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.16711