Fine-structure Preserved Real-world Image Super-resolution via Transfer VAE Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yi, Qiaosi, Li, Shuai, Wu, Rongyuan, Sun, Lingchen, Wu, Yuhui, Zhang, Lei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909708711886848
author Yi, Qiaosi
Li, Shuai
Wu, Rongyuan
Sun, Lingchen
Wu, Yuhui
Zhang, Lei
author_facet Yi, Qiaosi
Li, Shuai
Wu, Rongyuan
Sun, Lingchen
Wu, Yuhui
Zhang, Lei
contents Impressive results on real-world image super-resolution (Real-ISR) have been achieved by employing pre-trained stable diffusion (SD) models. However, one critical issue of such methods lies in their poor reconstruction of image fine structures, such as small characters and textures, due to the aggressive resolution reduction of the VAE (eg., 8$\times$ downsampling) in the SD model. One solution is to employ a VAE with a lower downsampling rate for diffusion; however, adapting its latent features with the pre-trained UNet while mitigating the increased computational cost poses new challenges. To address these issues, we propose a Transfer VAE Training (TVT) strategy to transfer the 8$\times$ downsampled VAE into a 4$\times$ one while adapting to the pre-trained UNet. Specifically, we first train a 4$\times$ decoder based on the output features of the original VAE encoder, then train a 4$\times$ encoder while keeping the newly trained decoder fixed. Such a TVT strategy aligns the new encoder-decoder pair with the original VAE latent space while enhancing image fine details. Additionally, we introduce a compact VAE and compute-efficient UNet by optimizing their network architectures, reducing the computational cost while capturing high-resolution fine-scale features. Experimental results demonstrate that our TVT method significantly improves fine-structure preservation, which is often compromised by other SD-based methods, while requiring fewer FLOPs than state-of-the-art one-step diffusion models. The official code can be found at https://github.com/Joyies/TVT.
format Preprint
id arxiv_https___arxiv_org_abs_2507_20291
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fine-structure Preserved Real-world Image Super-resolution via Transfer VAE Training
Yi, Qiaosi
Li, Shuai
Wu, Rongyuan
Sun, Lingchen
Wu, Yuhui
Zhang, Lei
Computer Vision and Pattern Recognition
Impressive results on real-world image super-resolution (Real-ISR) have been achieved by employing pre-trained stable diffusion (SD) models. However, one critical issue of such methods lies in their poor reconstruction of image fine structures, such as small characters and textures, due to the aggressive resolution reduction of the VAE (eg., 8$\times$ downsampling) in the SD model. One solution is to employ a VAE with a lower downsampling rate for diffusion; however, adapting its latent features with the pre-trained UNet while mitigating the increased computational cost poses new challenges. To address these issues, we propose a Transfer VAE Training (TVT) strategy to transfer the 8$\times$ downsampled VAE into a 4$\times$ one while adapting to the pre-trained UNet. Specifically, we first train a 4$\times$ decoder based on the output features of the original VAE encoder, then train a 4$\times$ encoder while keeping the newly trained decoder fixed. Such a TVT strategy aligns the new encoder-decoder pair with the original VAE latent space while enhancing image fine details. Additionally, we introduce a compact VAE and compute-efficient UNet by optimizing their network architectures, reducing the computational cost while capturing high-resolution fine-scale features. Experimental results demonstrate that our TVT method significantly improves fine-structure preservation, which is often compromised by other SD-based methods, while requiring fewer FLOPs than state-of-the-art one-step diffusion models. The official code can be found at https://github.com/Joyies/TVT.
title Fine-structure Preserved Real-world Image Super-resolution via Transfer VAE Training
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.20291