Training-Free Motion Customization for Distilled Video Generators with Adaptive Test-Time Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rong, Jintao, Xie, Xin, Yu, Xinyi, Ou, Linlin, Zhang, Xinyu, Shen, Chunhua, Gong, Dong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909658774503424
author Rong, Jintao
Xie, Xin
Yu, Xinyi
Ou, Linlin
Zhang, Xinyu
Shen, Chunhua
Gong, Dong
author_facet Rong, Jintao
Xie, Xin
Yu, Xinyi
Ou, Linlin
Zhang, Xinyu
Shen, Chunhua
Gong, Dong
contents Distilled video generation models offer fast and efficient synthesis but struggle with motion customization when guided by reference videos, especially under training-free settings. Existing training-free methods, originally designed for standard diffusion models, fail to generalize due to the accelerated generative process and large denoising steps in distilled models. To address this, we propose MotionEcho, a novel training-free test-time distillation framework that enables motion customization by leveraging diffusion teacher forcing. Our approach uses high-quality, slow teacher models to guide the inference of fast student models through endpoint prediction and interpolation. To maintain efficiency, we dynamically allocate computation across timesteps according to guidance needs. Extensive experiments across various distilled video generation models and benchmark datasets demonstrate that our method significantly improves motion fidelity and generation quality while preserving high efficiency. Project page: https://euminds.github.io/motionecho/
format Preprint
id arxiv_https___arxiv_org_abs_2506_19348
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Training-Free Motion Customization for Distilled Video Generators with Adaptive Test-Time Distillation
Rong, Jintao
Xie, Xin
Yu, Xinyi
Ou, Linlin
Zhang, Xinyu
Shen, Chunhua
Gong, Dong
Computer Vision and Pattern Recognition
Distilled video generation models offer fast and efficient synthesis but struggle with motion customization when guided by reference videos, especially under training-free settings. Existing training-free methods, originally designed for standard diffusion models, fail to generalize due to the accelerated generative process and large denoising steps in distilled models. To address this, we propose MotionEcho, a novel training-free test-time distillation framework that enables motion customization by leveraging diffusion teacher forcing. Our approach uses high-quality, slow teacher models to guide the inference of fast student models through endpoint prediction and interpolation. To maintain efficiency, we dynamically allocate computation across timesteps according to guidance needs. Extensive experiments across various distilled video generation models and benchmark datasets demonstrate that our method significantly improves motion fidelity and generation quality while preserving high efficiency. Project page: https://euminds.github.io/motionecho/
title Training-Free Motion Customization for Distilled Video Generators with Adaptive Test-Time Distillation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.19348