Structure From Tracking: Distilling Structure-Preserving Motion for Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Fei, Yang, Stoica, George, Liu, Jingyuan, Chen, Qifeng, Krishna, Ranjay, Wang, Xiaojuan, Liu, Benlin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual Representations inside the Language Model
by: Liu, Benlin, et al.
Published: (2025)
by: Liu, Benlin, et al.
Published: (2025)
Efficient Inference of Vision Instruction-Following Models with Elastic Cache
by: Liu, Zuyan, et al.
Published: (2024)
by: Liu, Zuyan, et al.
Published: (2024)
CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation
by: Su, Xia, et al.
Published: (2026)
by: Su, Xia, et al.
Published: (2026)
Contrastive Flow Matching
by: Stoica, George, et al.
Published: (2025)
by: Stoica, George, et al.
Published: (2025)
Large Motion Video Autoencoding with Cross-modal Video VAE
by: Xing, Yazhou, et al.
Published: (2024)
by: Xing, Yazhou, et al.
Published: (2024)
Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
by: Chen, Dongping, et al.
Published: (2024)
by: Chen, Dongping, et al.
Published: (2024)
Seeking and Updating with Live Visual Knowledge
by: Fu, Mingyang, et al.
Published: (2025)
by: Fu, Mingyang, et al.
Published: (2025)
RefTok: Reference-Based Tokenization for Video Generation
by: Fan, Xiang, et al.
Published: (2025)
by: Fan, Xiang, et al.
Published: (2025)
Motion Semantics Guided Normalizing Flow for Privacy-Preserving Video Anomaly Detection
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
Show and Polish: Reference-Guided Identity Preservation in Face Video Restoration
by: Han, Wenkang, et al.
Published: (2025)
by: Han, Wenkang, et al.
Published: (2025)
Learning Quantised Structure-Preserving Motion Representations for Dance Fingerprinting
by: Kharlamova, Arina, et al.
Published: (2026)
by: Kharlamova, Arina, et al.
Published: (2026)
RefDecoder: Enhancing Visual Generation with Conditional Video Decoding
by: Fan, Xiang, et al.
Published: (2026)
by: Fan, Xiang, et al.
Published: (2026)
Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model
by: Liu, Benlin, et al.
Published: (2024)
by: Liu, Benlin, et al.
Published: (2024)
Identity-Preserving Video Dubbing Using Motion Warping
by: Liu, Runzhen, et al.
Published: (2025)
by: Liu, Runzhen, et al.
Published: (2025)
Posterior Augmented Flow Matching
by: Stoica, George, et al.
Published: (2026)
by: Stoica, George, et al.
Published: (2026)
PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning
by: Li, Shaoxuan, et al.
Published: (2026)
by: Li, Shaoxuan, et al.
Published: (2026)
Streaming Autoregressive Video Generation via Diagonal Distillation
by: Liu, Jinxiu, et al.
Published: (2026)
by: Liu, Jinxiu, et al.
Published: (2026)
Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning
by: Ma, Yue, et al.
Published: (2025)
by: Ma, Yue, et al.
Published: (2025)
Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation
by: Zhai, Yuanhao, et al.
Published: (2024)
by: Zhai, Yuanhao, et al.
Published: (2024)
Generative Video Motion Editing with 3D Point Tracks
by: Lee, Yao-Chih, et al.
Published: (2025)
by: Lee, Yao-Chih, et al.
Published: (2025)
HeadArtist: Text-conditioned 3D Head Generation with Self Score Distillation
by: Liu, Hongyu, et al.
Published: (2023)
by: Liu, Hongyu, et al.
Published: (2023)
SpikeMOT: Event-based Multi-Object Tracking with Sparse Motion Features
by: Wang, Song, et al.
Published: (2023)
by: Wang, Song, et al.
Published: (2023)
Training-Free Motion Customization for Distilled Video Generators with Adaptive Test-Time Distillation
by: Rong, Jintao, et al.
Published: (2025)
by: Rong, Jintao, et al.
Published: (2025)
Veda: Scalable Video Diffusion via Distilled Sparse Attention
by: Han, Shihao, et al.
Published: (2026)
by: Han, Shihao, et al.
Published: (2026)
DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization
by: Ding, Zihan, et al.
Published: (2024)
by: Ding, Zihan, et al.
Published: (2024)
Zero-shot Synthetic Video Realism Enhancement via Structure-aware Denoising
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
by: Che, Liwei, et al.
Published: (2026)
by: Che, Liwei, et al.
Published: (2026)
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
by: Zhang, Tianyi, et al.
Published: (2026)
by: Zhang, Tianyi, et al.
Published: (2026)
Fast Video Generation with Sliding Tile Attention
by: Zhang, Peiyuan, et al.
Published: (2025)
by: Zhang, Peiyuan, et al.
Published: (2025)
TrackSSM: A General Motion Predictor by State-Space Model
by: Hu, Bin, et al.
Published: (2024)
by: Hu, Bin, et al.
Published: (2024)
FastVMT: Eliminating Redundancy in Video Motion Transfer
by: Ma, Yue, et al.
Published: (2026)
by: Ma, Yue, et al.
Published: (2026)
MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
Human Motion Video Generation: A Survey
by: Xue, Haiwei, et al.
Published: (2025)
by: Xue, Haiwei, et al.
Published: (2025)
Videoshop: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion
by: Fan, Xiang, et al.
Published: (2024)
by: Fan, Xiang, et al.
Published: (2024)
GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure
by: Gu, Leslie, et al.
Published: (2025)
by: Gu, Leslie, et al.
Published: (2025)
Consistency-Preserving Diverse Video Generation
by: Liu, Xinshuang, et al.
Published: (2026)
by: Liu, Xinshuang, et al.
Published: (2026)
SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
by: Brown, Ellis, et al.
Published: (2025)
by: Brown, Ellis, et al.
Published: (2025)
Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models
by: Hu, Yushi, et al.
Published: (2023)
by: Hu, Yushi, et al.
Published: (2023)
WildActor: Unconstrained Identity-Preserving Video Generation
by: Guo, Qin, et al.
Published: (2026)
by: Guo, Qin, et al.
Published: (2026)
Similar Items
-
Visual Representations inside the Language Model
by: Liu, Benlin, et al.
Published: (2025) -
Efficient Inference of Vision Instruction-Following Models with Elastic Cache
by: Liu, Zuyan, et al.
Published: (2024) -
CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation
by: Su, Xia, et al.
Published: (2026) -
Contrastive Flow Matching
by: Stoica, George, et al.
Published: (2025) -
Large Motion Video Autoencoding with Cross-modal Video VAE
by: Xing, Yazhou, et al.
Published: (2024)