DiC: Rethinking Conv3x3 Designs in Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tian, Yuchuan, Han, Jing, Wang, Chengcheng, Liang, Yuchen, Xu, Chao, Chen, Hanting
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908398302265344
author Tian, Yuchuan
Han, Jing
Wang, Chengcheng
Liang, Yuchen
Xu, Chao
Chen, Hanting
author_facet Tian, Yuchuan
Han, Jing
Wang, Chengcheng
Liang, Yuchen
Xu, Chao
Chen, Hanting
contents Diffusion models have shown exceptional performance in visual generation tasks. Recently, these models have shifted from traditional U-Shaped CNN-Attention hybrid structures to fully transformer-based isotropic architectures. While these transformers exhibit strong scalability and performance, their reliance on complicated self-attention operation results in slow inference speeds. Contrary to these works, we rethink one of the simplest yet fastest module in deep learning, 3x3 Convolution, to construct a scaled-up purely convolutional diffusion model. We first discover that an Encoder-Decoder Hourglass design outperforms scalable isotropic architectures for Conv3x3, but still under-performing our expectation. Further improving the architecture, we introduce sparse skip connections to reduce redundancy and improve scalability. Based on the architecture, we introduce conditioning improvements including stage-specific embeddings, mid-block condition injection, and conditional gating. These improvements lead to our proposed Diffusion CNN (DiC), which serves as a swift yet competitive diffusion architecture baseline. Experiments on various scales and settings show that DiC surpasses existing diffusion transformers by considerable margins in terms of performance while keeping a good speed advantage. Project page: https://github.com/YuchuanTian/DiC
format Preprint
id arxiv_https___arxiv_org_abs_2501_00603
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DiC: Rethinking Conv3x3 Designs in Diffusion Models
Tian, Yuchuan
Han, Jing
Wang, Chengcheng
Liang, Yuchen
Xu, Chao
Chen, Hanting
Computer Vision and Pattern Recognition
Machine Learning
Diffusion models have shown exceptional performance in visual generation tasks. Recently, these models have shifted from traditional U-Shaped CNN-Attention hybrid structures to fully transformer-based isotropic architectures. While these transformers exhibit strong scalability and performance, their reliance on complicated self-attention operation results in slow inference speeds. Contrary to these works, we rethink one of the simplest yet fastest module in deep learning, 3x3 Convolution, to construct a scaled-up purely convolutional diffusion model. We first discover that an Encoder-Decoder Hourglass design outperforms scalable isotropic architectures for Conv3x3, but still under-performing our expectation. Further improving the architecture, we introduce sparse skip connections to reduce redundancy and improve scalability. Based on the architecture, we introduce conditioning improvements including stage-specific embeddings, mid-block condition injection, and conditional gating. These improvements lead to our proposed Diffusion CNN (DiC), which serves as a swift yet competitive diffusion architecture baseline. Experiments on various scales and settings show that DiC surpasses existing diffusion transformers by considerable margins in terms of performance while keeping a good speed advantage. Project page: https://github.com/YuchuanTian/DiC
title DiC: Rethinking Conv3x3 Designs in Diffusion Models
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2501.00603