Structure From Tracking: Distilling Structure-Preserving Motion for Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fei, Yang, Stoica, George, Liu, Jingyuan, Chen, Qifeng, Krishna, Ranjay, Wang, Xiaojuan, Liu, Benlin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917142812688384
author Fei, Yang
Stoica, George
Liu, Jingyuan
Chen, Qifeng
Krishna, Ranjay
Wang, Xiaojuan
Liu, Benlin
author_facet Fei, Yang
Stoica, George
Liu, Jingyuan
Chen, Qifeng
Krishna, Ranjay
Wang, Xiaojuan
Liu, Benlin
contents Reality is a dance between rigid constraints and deformable structures. For video models, that means generating motion that preserves fidelity as well as structure. Despite progress in diffusion models, producing realistic structure-preserving motion remains challenging, especially for articulated and deformable objects such as humans and animals. Scaling training data alone, so far, has failed to resolve physically implausible transitions. Existing approaches rely on conditioning with noisy motion representations, such as optical flow or skeletons extracted using an external imperfect model. To address these challenges, we introduce an algorithm to distill structure-preserving motion priors from an autoregressive video tracking model (SAM2) into a bidirectional video diffusion model (CogVideoX). With our method, we train SAM2VideoX, which contains two innovations: (1) a bidirectional feature fusion module that extracts global structure-preserving motion priors from a recurrent model like SAM2; (2) a Local Gram Flow loss that aligns how local features move together. Experiments on VBench and in human studies show that SAM2VideoX delivers consistent gains (+2.60\% on VBench, 21-22\% lower FVD, and 71.4\% human preference) over prior baselines. Specifically, on VBench, we achieve 95.51\%, surpassing REPA (92.91\%) by 2.60\%, and reduce FVD to 360.57, a 21.20\% and 22.46\% improvement over REPA- and LoRA-finetuning, respectively. The project website can be found at https://sam2videox.github.io/ .
format Preprint
id arxiv_https___arxiv_org_abs_2512_11792
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Structure From Tracking: Distilling Structure-Preserving Motion for Video Generation
Fei, Yang
Stoica, George
Liu, Jingyuan
Chen, Qifeng
Krishna, Ranjay
Wang, Xiaojuan
Liu, Benlin
Computer Vision and Pattern Recognition
Reality is a dance between rigid constraints and deformable structures. For video models, that means generating motion that preserves fidelity as well as structure. Despite progress in diffusion models, producing realistic structure-preserving motion remains challenging, especially for articulated and deformable objects such as humans and animals. Scaling training data alone, so far, has failed to resolve physically implausible transitions. Existing approaches rely on conditioning with noisy motion representations, such as optical flow or skeletons extracted using an external imperfect model. To address these challenges, we introduce an algorithm to distill structure-preserving motion priors from an autoregressive video tracking model (SAM2) into a bidirectional video diffusion model (CogVideoX). With our method, we train SAM2VideoX, which contains two innovations: (1) a bidirectional feature fusion module that extracts global structure-preserving motion priors from a recurrent model like SAM2; (2) a Local Gram Flow loss that aligns how local features move together. Experiments on VBench and in human studies show that SAM2VideoX delivers consistent gains (+2.60\% on VBench, 21-22\% lower FVD, and 71.4\% human preference) over prior baselines. Specifically, on VBench, we achieve 95.51\%, surpassing REPA (92.91\%) by 2.60\%, and reduce FVD to 360.57, a 21.20\% and 22.46\% improvement over REPA- and LoRA-finetuning, respectively. The project website can be found at https://sam2videox.github.io/ .
title Structure From Tracking: Distilling Structure-Preserving Motion for Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.11792