Training-Free Motion Customization for Distilled Video Generators with Adaptive Test-Time Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909658774503424 |
|---|---|
| author | Rong, Jintao Xie, Xin Yu, Xinyi Ou, Linlin Zhang, Xinyu Shen, Chunhua Gong, Dong |
| author_facet | Rong, Jintao Xie, Xin Yu, Xinyi Ou, Linlin Zhang, Xinyu Shen, Chunhua Gong, Dong |
| contents | Distilled video generation models offer fast and efficient synthesis but struggle with motion customization when guided by reference videos, especially under training-free settings. Existing training-free methods, originally designed for standard diffusion models, fail to generalize due to the accelerated generative process and large denoising steps in distilled models. To address this, we propose MotionEcho, a novel training-free test-time distillation framework that enables motion customization by leveraging diffusion teacher forcing. Our approach uses high-quality, slow teacher models to guide the inference of fast student models through endpoint prediction and interpolation. To maintain efficiency, we dynamically allocate computation across timesteps according to guidance needs. Extensive experiments across various distilled video generation models and benchmark datasets demonstrate that our method significantly improves motion fidelity and generation quality while preserving high efficiency. Project page: https://euminds.github.io/motionecho/ |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_19348 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Training-Free Motion Customization for Distilled Video Generators with Adaptive Test-Time Distillation Rong, Jintao Xie, Xin Yu, Xinyi Ou, Linlin Zhang, Xinyu Shen, Chunhua Gong, Dong Computer Vision and Pattern Recognition Distilled video generation models offer fast and efficient synthesis but struggle with motion customization when guided by reference videos, especially under training-free settings. Existing training-free methods, originally designed for standard diffusion models, fail to generalize due to the accelerated generative process and large denoising steps in distilled models. To address this, we propose MotionEcho, a novel training-free test-time distillation framework that enables motion customization by leveraging diffusion teacher forcing. Our approach uses high-quality, slow teacher models to guide the inference of fast student models through endpoint prediction and interpolation. To maintain efficiency, we dynamically allocate computation across timesteps according to guidance needs. Extensive experiments across various distilled video generation models and benchmark datasets demonstrate that our method significantly improves motion fidelity and generation quality while preserving high efficiency. Project page: https://euminds.github.io/motionecho/ |
| title | Training-Free Motion Customization for Distilled Video Generators with Adaptive Test-Time Distillation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2506.19348 |