Asymmetric VAE for One-Step Video Super-Resolution Acceleration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jianze, Guo, Yong, Zhang, Yulun, Yang, Xiaokang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909813353480192
author Li, Jianze
Guo, Yong
Zhang, Yulun
Yang, Xiaokang
author_facet Li, Jianze
Guo, Yong
Zhang, Yulun
Yang, Xiaokang
contents Diffusion models have significant advantages in the field of real-world video super-resolution and have demonstrated strong performance in past research. In recent diffusion-based video super-resolution (VSR) models, the number of sampling steps has been reduced to just one, yet there remains significant room for further optimization in inference efficiency. In this paper, we propose FastVSR, which achieves substantial reductions in computational cost by implementing a high compression VAE (spatial compression ratio of 16, denoted as f16). We design the structure of the f16 VAE and introduce a stable training framework. We employ pixel shuffle and channel replication to achieve additional upsampling. Furthermore, we propose a lower-bound-guided training strategy, which introduces a simpler training objective as a lower bound for the VAE's performance. It makes the training process more stable and easier to converge. Experimental results show that FastVSR achieves speedups of 111.9 times compared to multi-step models and 3.92 times compared to existing one-step models. We will release code and models at https://github.com/JianzeLi-114/FastVSR.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24142
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Asymmetric VAE for One-Step Video Super-Resolution Acceleration
Li, Jianze
Guo, Yong
Zhang, Yulun
Yang, Xiaokang
Computer Vision and Pattern Recognition
Diffusion models have significant advantages in the field of real-world video super-resolution and have demonstrated strong performance in past research. In recent diffusion-based video super-resolution (VSR) models, the number of sampling steps has been reduced to just one, yet there remains significant room for further optimization in inference efficiency. In this paper, we propose FastVSR, which achieves substantial reductions in computational cost by implementing a high compression VAE (spatial compression ratio of 16, denoted as f16). We design the structure of the f16 VAE and introduce a stable training framework. We employ pixel shuffle and channel replication to achieve additional upsampling. Furthermore, we propose a lower-bound-guided training strategy, which introduces a simpler training objective as a lower bound for the VAE's performance. It makes the training process more stable and easier to converge. Experimental results show that FastVSR achieves speedups of 111.9 times compared to multi-step models and 3.92 times compared to existing one-step models. We will release code and models at https://github.com/JianzeLi-114/FastVSR.
title Asymmetric VAE for One-Step Video Super-Resolution Acceleration
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.24142