ControlNeXt: Powerful and Efficient Control for Image and Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peng, Bohao, Wang, Jian, Zhang, Yuechen, Li, Wenbo, Yang, Ming-Chang, Jia, Jiaya
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912265184215040
author Peng, Bohao
Wang, Jian
Zhang, Yuechen
Li, Wenbo
Yang, Ming-Chang
Jia, Jiaya
author_facet Peng, Bohao
Wang, Jian
Zhang, Yuechen
Li, Wenbo
Yang, Ming-Chang
Jia, Jiaya
contents Diffusion models have demonstrated remarkable and robust abilities in both image and video generation. To achieve greater control over generated results, researchers introduce additional architectures, such as ControlNet, Adapters and ReferenceNet, to integrate conditioning controls. However, current controllable generation methods often require substantial additional computational resources, especially for video generation, and face challenges in training or exhibit weak control. In this paper, we propose ControlNeXt: a powerful and efficient method for controllable image and video generation. We first design a more straightforward and efficient architecture, replacing heavy additional branches with minimal additional cost compared to the base model. Such a concise structure also allows our method to seamlessly integrate with other LoRA weights, enabling style alteration without the need for additional training. As for training, we reduce up to 90% of learnable parameters compared to the alternatives. Furthermore, we propose another method called Cross Normalization (CN) as a replacement for Zero-Convolution' to achieve fast and stable training convergence. We have conducted various experiments with different base models across images and videos, demonstrating the robustness of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2408_06070
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ControlNeXt: Powerful and Efficient Control for Image and Video Generation
Peng, Bohao
Wang, Jian
Zhang, Yuechen
Li, Wenbo
Yang, Ming-Chang
Jia, Jiaya
Computer Vision and Pattern Recognition
Diffusion models have demonstrated remarkable and robust abilities in both image and video generation. To achieve greater control over generated results, researchers introduce additional architectures, such as ControlNet, Adapters and ReferenceNet, to integrate conditioning controls. However, current controllable generation methods often require substantial additional computational resources, especially for video generation, and face challenges in training or exhibit weak control. In this paper, we propose ControlNeXt: a powerful and efficient method for controllable image and video generation. We first design a more straightforward and efficient architecture, replacing heavy additional branches with minimal additional cost compared to the base model. Such a concise structure also allows our method to seamlessly integrate with other LoRA weights, enabling style alteration without the need for additional training. As for training, we reduce up to 90% of learnable parameters compared to the alternatives. Furthermore, we propose another method called Cross Normalization (CN) as a replacement for Zero-Convolution' to achieve fast and stable training convergence. We have conducted various experiments with different base models across images and videos, demonstrating the robustness of our method.
title ControlNeXt: Powerful and Efficient Control for Image and Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.06070