FreeScale: Unleashing the Resolution of Diffusion Models via Tuning-Free Scale Fusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qiu, Haonan, Zhang, Shiwei, Wei, Yujie, Chu, Ruihang, Yuan, Hangjie, Wang, Xiang, Zhang, Yingya, Liu, Ziwei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909683369902080
author Qiu, Haonan
Zhang, Shiwei
Wei, Yujie
Chu, Ruihang
Yuan, Hangjie
Wang, Xiang
Zhang, Yingya
Liu, Ziwei
author_facet Qiu, Haonan
Zhang, Shiwei
Wei, Yujie
Chu, Ruihang
Yuan, Hangjie
Wang, Xiang
Zhang, Yingya
Liu, Ziwei
contents Visual diffusion models achieve remarkable progress, yet they are typically trained at limited resolutions due to the lack of high-resolution data and constrained computation resources, hampering their ability to generate high-fidelity images or videos at higher resolutions. Recent efforts have explored tuning-free strategies to exhibit the untapped potential higher-resolution visual generation of pre-trained models. However, these methods are still prone to producing low-quality visual content with repetitive patterns. The key obstacle lies in the inevitable increase in high-frequency information when the model generates visual content exceeding its training resolution, leading to undesirable repetitive patterns deriving from the accumulated errors. To tackle this challenge, we propose FreeScale, a tuning-free inference paradigm to enable higher-resolution visual generation via scale fusion. Specifically, FreeScale processes information from different receptive scales and then fuses it by extracting desired frequency components. Extensive experiments validate the superiority of our paradigm in extending the capabilities of higher-resolution visual generation for both image and video models. Notably, compared with previous best-performing methods, FreeScale unlocks the 8k-resolution text-to-image generation for the first time.
format Preprint
id arxiv_https___arxiv_org_abs_2412_09626
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FreeScale: Unleashing the Resolution of Diffusion Models via Tuning-Free Scale Fusion
Qiu, Haonan
Zhang, Shiwei
Wei, Yujie
Chu, Ruihang
Yuan, Hangjie
Wang, Xiang
Zhang, Yingya
Liu, Ziwei
Computer Vision and Pattern Recognition
Visual diffusion models achieve remarkable progress, yet they are typically trained at limited resolutions due to the lack of high-resolution data and constrained computation resources, hampering their ability to generate high-fidelity images or videos at higher resolutions. Recent efforts have explored tuning-free strategies to exhibit the untapped potential higher-resolution visual generation of pre-trained models. However, these methods are still prone to producing low-quality visual content with repetitive patterns. The key obstacle lies in the inevitable increase in high-frequency information when the model generates visual content exceeding its training resolution, leading to undesirable repetitive patterns deriving from the accumulated errors. To tackle this challenge, we propose FreeScale, a tuning-free inference paradigm to enable higher-resolution visual generation via scale fusion. Specifically, FreeScale processes information from different receptive scales and then fuses it by extracting desired frequency components. Extensive experiments validate the superiority of our paradigm in extending the capabilities of higher-resolution visual generation for both image and video models. Notably, compared with previous best-performing methods, FreeScale unlocks the 8k-resolution text-to-image generation for the first time.
title FreeScale: Unleashing the Resolution of Diffusion Models via Tuning-Free Scale Fusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.09626