UFO: Enhancing Diffusion-Based Video Generation with a Uniform Frame Organizer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Delong, Hou, Zhaohui, Zhan, Mingjie, Han, Shihao, Zhao, Zhicheng, Su, Fei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910743255842816
author Liu, Delong
Hou, Zhaohui
Zhan, Mingjie
Han, Shihao
Zhao, Zhicheng
Su, Fei
author_facet Liu, Delong
Hou, Zhaohui
Zhan, Mingjie
Han, Shihao
Zhao, Zhicheng
Su, Fei
contents Recently, diffusion-based video generation models have achieved significant success. However, existing models often suffer from issues like weak consistency and declining image quality over time. To overcome these challenges, inspired by aesthetic principles, we propose a non-invasive plug-in called Uniform Frame Organizer (UFO), which is compatible with any diffusion-based video generation model. The UFO comprises a series of adaptive adapters with adjustable intensities, which can significantly enhance the consistency between the foreground and background of videos and improve image quality without altering the original model parameters when integrated. The training for UFO is simple, efficient, requires minimal resources, and supports stylized training. Its modular design allows for the combination of multiple UFOs, enabling the customization of personalized video generation models. Furthermore, the UFO also supports direct transferability across different models of the same specification without the need for specific retraining. The experimental results indicate that UFO effectively enhances video generation quality and demonstrates its superiority in public video generation benchmarks. The code will be publicly available at https://github.com/Delong-liu-bupt/UFO.
format Preprint
id arxiv_https___arxiv_org_abs_2412_09389
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle UFO: Enhancing Diffusion-Based Video Generation with a Uniform Frame Organizer
Liu, Delong
Hou, Zhaohui
Zhan, Mingjie
Han, Shihao
Zhao, Zhicheng
Su, Fei
Computer Vision and Pattern Recognition
Artificial Intelligence
Recently, diffusion-based video generation models have achieved significant success. However, existing models often suffer from issues like weak consistency and declining image quality over time. To overcome these challenges, inspired by aesthetic principles, we propose a non-invasive plug-in called Uniform Frame Organizer (UFO), which is compatible with any diffusion-based video generation model. The UFO comprises a series of adaptive adapters with adjustable intensities, which can significantly enhance the consistency between the foreground and background of videos and improve image quality without altering the original model parameters when integrated. The training for UFO is simple, efficient, requires minimal resources, and supports stylized training. Its modular design allows for the combination of multiple UFOs, enabling the customization of personalized video generation models. Furthermore, the UFO also supports direct transferability across different models of the same specification without the need for specific retraining. The experimental results indicate that UFO effectively enhances video generation quality and demonstrates its superiority in public video generation benchmarks. The code will be publicly available at https://github.com/Delong-liu-bupt/UFO.
title UFO: Enhancing Diffusion-Based Video Generation with a Uniform Frame Organizer
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2412.09389