SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Xinyao, Dong, Wenkai, Song, Yuxin, Fang, Bo, Zhang, Qi, Wang, Jing, Chen, Fan, Zhang, Hui, Feng, Haocheng, Lu, Yu, Zhou, Hang, Yuan, Chun, Wang, Jingdong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911530588569600
author Zhang, Xinyao
Dong, Wenkai
Song, Yuxin
Fang, Bo
Zhang, Qi
Wang, Jing
Chen, Fan
Zhang, Hui
Feng, Haocheng
Lu, Yu
Zhou, Hang
Yuan, Chun
Wang, Jingdong
author_facet Zhang, Xinyao
Dong, Wenkai
Song, Yuxin
Fang, Bo
Zhang, Qi
Wang, Jing
Chen, Fan
Zhang, Hui
Feng, Haocheng
Lu, Yu
Zhou, Hang
Yuan, Chun
Wang, Jingdong
contents Current instruction-guided video editing models struggle to simultaneously balance precise semantic modifications with faithful motion preservation. While existing approaches rely on injecting explicit external priors (e.g., VLM features or structural conditions) to mitigate these issues, this reliance severely bottlenecks model robustness and generalization. To overcome this limitation, we present SAMA (factorized Semantic Anchoring and Motion Alignment), a framework that factorizes video editing into semantic anchoring and motion modeling. First, we introduce Semantic Anchoring, which establishes a reliable visual anchor by jointly predicting semantic tokens and video latents at sparse anchor frames, enabling purely instruction-aware structural planning. Second, Motion Alignment pre-trains the same backbone on motion-centric video restoration pretext tasks (cube inpainting, speed perturbation, and tube shuffle), enabling the model to internalize temporal dynamics directly from raw videos. SAMA is optimized with a two-stage pipeline: a factorized pre-training stage that learns inherent semantic-motion representations without paired video-instruction editing data, followed by supervised fine-tuning on paired editing data. Remarkably, the factorized pre-training alone already yields strong zero-shot video editing ability, validating the proposed factorization. SAMA achieves state-of-the-art performance among open-source models and is competitive with leading commercial systems (e.g., Kling-Omni). Code, models, and datasets will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2603_19228
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing
Zhang, Xinyao
Dong, Wenkai
Song, Yuxin
Fang, Bo
Zhang, Qi
Wang, Jing
Chen, Fan
Zhang, Hui
Feng, Haocheng
Lu, Yu
Zhou, Hang
Yuan, Chun
Wang, Jingdong
Computer Vision and Pattern Recognition
Current instruction-guided video editing models struggle to simultaneously balance precise semantic modifications with faithful motion preservation. While existing approaches rely on injecting explicit external priors (e.g., VLM features or structural conditions) to mitigate these issues, this reliance severely bottlenecks model robustness and generalization. To overcome this limitation, we present SAMA (factorized Semantic Anchoring and Motion Alignment), a framework that factorizes video editing into semantic anchoring and motion modeling. First, we introduce Semantic Anchoring, which establishes a reliable visual anchor by jointly predicting semantic tokens and video latents at sparse anchor frames, enabling purely instruction-aware structural planning. Second, Motion Alignment pre-trains the same backbone on motion-centric video restoration pretext tasks (cube inpainting, speed perturbation, and tube shuffle), enabling the model to internalize temporal dynamics directly from raw videos. SAMA is optimized with a two-stage pipeline: a factorized pre-training stage that learns inherent semantic-motion representations without paired video-instruction editing data, followed by supervised fine-tuning on paired editing data. Remarkably, the factorized pre-training alone already yields strong zero-shot video editing ability, validating the proposed factorization. SAMA achieves state-of-the-art performance among open-source models and is competitive with leading commercial systems (e.g., Kling-Omni). Code, models, and datasets will be released.
title SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.19228