SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Yushu, Zhang, Zhixing, Li, Yanyu, Xu, Yanwu, Kag, Anil, Sui, Yang, Coskun, Huseyin, Ma, Ke, Lebedev, Aleksei, Hu, Ju, Metaxas, Dimitris, Wang, Yanzhi, Tulyakov, Sergey, Ren, Jian
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912421342347264
author Wu, Yushu
Zhang, Zhixing
Li, Yanyu
Xu, Yanwu
Kag, Anil
Sui, Yang
Coskun, Huseyin
Ma, Ke
Lebedev, Aleksei
Hu, Ju
Metaxas, Dimitris
Wang, Yanzhi
Tulyakov, Sergey
Ren, Jian
author_facet Wu, Yushu
Zhang, Zhixing
Li, Yanyu
Xu, Yanwu
Kag, Anil
Sui, Yang
Coskun, Huseyin
Ma, Ke
Lebedev, Aleksei
Hu, Ju
Metaxas, Dimitris
Wang, Yanzhi
Tulyakov, Sergey
Ren, Jian
contents We have witnessed the unprecedented success of diffusion-based video generation over the past year. Recently proposed models from the community have wielded the power to generate cinematic and high-resolution videos with smooth motions from arbitrary input prompts. However, as a supertask of image generation, video generation models require more computation and are thus hosted mostly on cloud servers, limiting broader adoption among content creators. In this work, we propose a comprehensive acceleration framework to bring the power of the large-scale video diffusion model to the hands of edge users. From the network architecture scope, we initialize from a compact image backbone and search out the design and arrangement of temporal layers to maximize hardware efficiency. In addition, we propose a dedicated adversarial fine-tuning algorithm for our efficient model and reduce the denoising steps to 4. Our model, with only 0.6B parameters, can generate a 5-second video on an iPhone 16 PM within 5 seconds. Compared to server-side models that take minutes on powerful GPUs to generate a single video, we accelerate the generation by magnitudes while delivering on-par quality.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10494
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device
Wu, Yushu
Zhang, Zhixing
Li, Yanyu
Xu, Yanwu
Kag, Anil
Sui, Yang
Coskun, Huseyin
Ma, Ke
Lebedev, Aleksei
Hu, Ju
Metaxas, Dimitris
Wang, Yanzhi
Tulyakov, Sergey
Ren, Jian
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Performance
We have witnessed the unprecedented success of diffusion-based video generation over the past year. Recently proposed models from the community have wielded the power to generate cinematic and high-resolution videos with smooth motions from arbitrary input prompts. However, as a supertask of image generation, video generation models require more computation and are thus hosted mostly on cloud servers, limiting broader adoption among content creators. In this work, we propose a comprehensive acceleration framework to bring the power of the large-scale video diffusion model to the hands of edge users. From the network architecture scope, we initialize from a compact image backbone and search out the design and arrangement of temporal layers to maximize hardware efficiency. In addition, we propose a dedicated adversarial fine-tuning algorithm for our efficient model and reduce the denoising steps to 4. Our model, with only 0.6B parameters, can generate a 5-second video on an iPhone 16 PM within 5 seconds. Compared to server-side models that take minutes on powerful GPUs to generate a single video, we accelerate the generation by magnitudes while delivering on-par quality.
title SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Performance
url https://arxiv.org/abs/2412.10494