Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Fei, Zhengcong, Fan, Mingyuan, Yu, Changqian, Li, Debang, Huang, Junshi
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911828907393024
author Fei, Zhengcong
Fan, Mingyuan
Yu, Changqian
Li, Debang
Huang, Junshi
author_facet Fei, Zhengcong
Fan, Mingyuan
Yu, Changqian
Li, Debang
Huang, Junshi
contents Transformers have catalyzed advancements in computer vision and natural language processing (NLP) fields. However, substantial computational complexity poses limitations for their application in long-context tasks, such as high-resolution image generation. This paper introduces a series of architectures adapted from the RWKV model used in the NLP, with requisite modifications tailored for diffusion model applied to image generation tasks, referred to as Diffusion-RWKV. Similar to the diffusion with Transformers, our model is designed to efficiently handle patchnified inputs in a sequence with extra conditions, while also scaling up effectively, accommodating both large-scale parameters and extensive datasets. Its distinctive advantage manifests in its reduced spatial aggregation complexity, rendering it exceptionally adept at processing high-resolution images, thereby eliminating the necessity for windowing or group cached operations. Experimental results on both condition and unconditional image generation tasks demonstrate that Diffison-RWKV achieves performance on par with or surpasses existing CNN or Transformer-based diffusion models in FID and IS metrics while significantly reducing total computation FLOP usage.
format Preprint
id arxiv_https___arxiv_org_abs_2404_04478
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models
Fei, Zhengcong
Fan, Mingyuan
Yu, Changqian
Li, Debang
Huang, Junshi
Computer Vision and Pattern Recognition
Transformers have catalyzed advancements in computer vision and natural language processing (NLP) fields. However, substantial computational complexity poses limitations for their application in long-context tasks, such as high-resolution image generation. This paper introduces a series of architectures adapted from the RWKV model used in the NLP, with requisite modifications tailored for diffusion model applied to image generation tasks, referred to as Diffusion-RWKV. Similar to the diffusion with Transformers, our model is designed to efficiently handle patchnified inputs in a sequence with extra conditions, while also scaling up effectively, accommodating both large-scale parameters and extensive datasets. Its distinctive advantage manifests in its reduced spatial aggregation complexity, rendering it exceptionally adept at processing high-resolution images, thereby eliminating the necessity for windowing or group cached operations. Experimental results on both condition and unconditional image generation tasks demonstrate that Diffison-RWKV achieves performance on par with or surpasses existing CNN or Transformer-based diffusion models in FID and IS metrics while significantly reducing total computation FLOP usage.
title Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.04478