OS-DiffVSR: Towards One-step Latent Diffusion Model for High-detailed Real-world Video Super-Resolution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Hanting, Tang, Huaao, Han, Jianhong, Zhou, Tianxiong, Cui, Jiulong, Xie, Haizhen, Chen, Yan, Hu, Jie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914048713424896
author Li, Hanting
Tang, Huaao
Han, Jianhong
Zhou, Tianxiong
Cui, Jiulong
Xie, Haizhen
Chen, Yan
Hu, Jie
author_facet Li, Hanting
Tang, Huaao
Han, Jianhong
Zhou, Tianxiong
Cui, Jiulong
Xie, Haizhen
Chen, Yan
Hu, Jie
contents Recently, latent diffusion models has demonstrated promising performance in real-world video super-resolution (VSR) task, which can reconstruct high-quality videos from distorted low-resolution input through multiple diffusion steps. Compared to image super-resolution (ISR), VSR methods needs to process each frame in a video, which poses challenges to its inference efficiency. However, video quality and inference efficiency have always been a trade-off for the diffusion-based VSR methods. In this work, we propose One-Step Diffusion model for real-world Video Super-Resolution, namely OS-DiffVSR. Specifically, we devise a novel adjacent frame adversarial training paradigm, which can significantly improve the quality of synthetic videos. Besides, we devise a multi-frame fusion mechanism to maintain inter-frame temporal consistency and reduce the flicker in video. Extensive experiments on several popular VSR benchmarks demonstrate that OS-DiffVSR can even achieve better quality than existing diffusion-based VSR methods that require dozens of sampling steps.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16507
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OS-DiffVSR: Towards One-step Latent Diffusion Model for High-detailed Real-world Video Super-Resolution
Li, Hanting
Tang, Huaao
Han, Jianhong
Zhou, Tianxiong
Cui, Jiulong
Xie, Haizhen
Chen, Yan
Hu, Jie
Computer Vision and Pattern Recognition
Recently, latent diffusion models has demonstrated promising performance in real-world video super-resolution (VSR) task, which can reconstruct high-quality videos from distorted low-resolution input through multiple diffusion steps. Compared to image super-resolution (ISR), VSR methods needs to process each frame in a video, which poses challenges to its inference efficiency. However, video quality and inference efficiency have always been a trade-off for the diffusion-based VSR methods. In this work, we propose One-Step Diffusion model for real-world Video Super-Resolution, namely OS-DiffVSR. Specifically, we devise a novel adjacent frame adversarial training paradigm, which can significantly improve the quality of synthetic videos. Besides, we devise a multi-frame fusion mechanism to maintain inter-frame temporal consistency and reduce the flicker in video. Extensive experiments on several popular VSR benchmarks demonstrate that OS-DiffVSR can even achieve better quality than existing diffusion-based VSR methods that require dozens of sampling steps.
title OS-DiffVSR: Towards One-step Latent Diffusion Model for High-detailed Real-world Video Super-Resolution
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.16507