InstaVSR: Taming Diffusion for Efficient and Temporally Consistent Video Super-Resolution
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914426492289024 |
|---|---|
| author | Hu, Jintong Chen, Bin Hu, Zhenyu Liu, Jiayue Wang, Guo Qi, Lu |
| author_facet | Hu, Jintong Chen, Bin Hu, Zhenyu Liu, Jiayue Wang, Guo Qi, Lu |
| contents | Video super-resolution (VSR) seeks to reconstruct high-resolution frames from low-resolution inputs. While diffusion-based methods have substantially improved perceptual quality, extending them to video remains challenging for two reasons: strong generative priors can introduce temporal instability, and multi-frame diffusion pipelines are often too expensive for practical deployment. To address both challenges simultaneously, we propose InstaVSR, a lightweight diffusion framework for efficient video super-resolution. InstaVSR combines three ingredients: (1) a pruned one-step diffusion backbone that removes several costly components from conventional diffusion-based VSR pipelines, (2) recurrent training with flow-guided temporal regularization to improve frame-to-frame stability, and (3) dual-space adversarial learning in latent and pixel spaces to preserve perceptual quality after backbone simplification. On an NVIDIA RTX 4090, InstaVSR processes a 30-frame video at 2K$\times$2K resolution in under one minute with only 7 GB of memory usage, substantially reducing the computational cost compared to existing diffusion-based methods while maintaining favorable perceptual quality with significantly smoother temporal transitions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_26134 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | InstaVSR: Taming Diffusion for Efficient and Temporally Consistent Video Super-Resolution Hu, Jintong Chen, Bin Hu, Zhenyu Liu, Jiayue Wang, Guo Qi, Lu Computer Vision and Pattern Recognition Video super-resolution (VSR) seeks to reconstruct high-resolution frames from low-resolution inputs. While diffusion-based methods have substantially improved perceptual quality, extending them to video remains challenging for two reasons: strong generative priors can introduce temporal instability, and multi-frame diffusion pipelines are often too expensive for practical deployment. To address both challenges simultaneously, we propose InstaVSR, a lightweight diffusion framework for efficient video super-resolution. InstaVSR combines three ingredients: (1) a pruned one-step diffusion backbone that removes several costly components from conventional diffusion-based VSR pipelines, (2) recurrent training with flow-guided temporal regularization to improve frame-to-frame stability, and (3) dual-space adversarial learning in latent and pixel spaces to preserve perceptual quality after backbone simplification. On an NVIDIA RTX 4090, InstaVSR processes a 30-frame video at 2K$\times$2K resolution in under one minute with only 7 GB of memory usage, substantially reducing the computational cost compared to existing diffusion-based methods while maintaining favorable perceptual quality with significantly smoother temporal transitions. |
| title | InstaVSR: Taming Diffusion for Efficient and Temporally Consistent Video Super-Resolution |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2603.26134 |