Saved in:
Bibliographic Details
Main Authors: Yi, Hongwei, Shao, Shitong, Ye, Tian, Zhao, Jiantong, Yin, Qingyu, Lingelbach, Michael, Yuan, Li, Tian, Yonghong, Xie, Enze, Zhou, Daquan
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2502.07701
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913692173467648
author Yi, Hongwei
Shao, Shitong
Ye, Tian
Zhao, Jiantong
Yin, Qingyu
Lingelbach, Michael
Yuan, Li
Tian, Yonghong
Xie, Enze
Zhou, Daquan
author_facet Yi, Hongwei
Shao, Shitong
Ye, Tian
Zhao, Jiantong
Yin, Qingyu
Lingelbach, Michael
Yuan, Li
Tian, Yonghong
Xie, Enze
Zhou, Daquan
contents In this technical report, we present Magic 1-For-1 (Magic141), an efficient video generation model with optimized memory consumption and inference latency. The key idea is simple: factorize the text-to-video generation task into two separate easier tasks for diffusion step distillation, namely text-to-image generation and image-to-video generation. We verify that with the same optimization algorithm, the image-to-video task is indeed easier to converge over the text-to-video task. We also explore a bag of optimization tricks to reduce the computational cost of training the image-to-video (I2V) models from three aspects: 1) model convergence speedup by using a multi-modal prior condition injection; 2) inference latency speed up by applying an adversarial step distillation, and 3) inference memory cost optimization with parameter sparsification. With those techniques, we are able to generate 5-second video clips within 3 seconds. By applying a test time sliding window, we are able to generate a minute-long video within one minute with significantly improved visual quality and motion dynamics, spending less than 1 second for generating 1 second video clips on average. We conduct a series of preliminary explorations to find out the optimal tradeoff between computational cost and video quality during diffusion step distillation and hope this could be a good foundation model for open-source explorations. The code and the model weights are available at https://github.com/DA-Group-PKU/Magic-1-For-1.
format Preprint
id arxiv_https___arxiv_org_abs_2502_07701
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Magic 1-For-1: Generating One Minute Video Clips within One Minute
Yi, Hongwei
Shao, Shitong
Ye, Tian
Zhao, Jiantong
Yin, Qingyu
Lingelbach, Michael
Yuan, Li
Tian, Yonghong
Xie, Enze
Zhou, Daquan
Computer Vision and Pattern Recognition
In this technical report, we present Magic 1-For-1 (Magic141), an efficient video generation model with optimized memory consumption and inference latency. The key idea is simple: factorize the text-to-video generation task into two separate easier tasks for diffusion step distillation, namely text-to-image generation and image-to-video generation. We verify that with the same optimization algorithm, the image-to-video task is indeed easier to converge over the text-to-video task. We also explore a bag of optimization tricks to reduce the computational cost of training the image-to-video (I2V) models from three aspects: 1) model convergence speedup by using a multi-modal prior condition injection; 2) inference latency speed up by applying an adversarial step distillation, and 3) inference memory cost optimization with parameter sparsification. With those techniques, we are able to generate 5-second video clips within 3 seconds. By applying a test time sliding window, we are able to generate a minute-long video within one minute with significantly improved visual quality and motion dynamics, spending less than 1 second for generating 1 second video clips on average. We conduct a series of preliminary explorations to find out the optimal tradeoff between computational cost and video quality during diffusion step distillation and hope this could be a good foundation model for open-source explorations. The code and the model weights are available at https://github.com/DA-Group-PKU/Magic-1-For-1.
title Magic 1-For-1: Generating One Minute Video Clips within One Minute
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.07701