SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Jianyi, Lin, Zhijie, Wei, Meng, Zhao, Yang, Yang, Ceyuan, Xiao, Fei, Loy, Chen Change, Jiang, Lu
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915208858959872
author Wang, Jianyi
Lin, Zhijie
Wei, Meng
Zhao, Yang
Yang, Ceyuan
Xiao, Fei
Loy, Chen Change
Jiang, Lu
author_facet Wang, Jianyi
Lin, Zhijie
Wei, Meng
Zhao, Yang
Yang, Ceyuan
Xiao, Fei
Loy, Chen Change
Jiang, Lu
contents Video restoration poses non-trivial challenges in maintaining fidelity while recovering temporally consistent details from unknown degradations in the wild. Despite recent advances in diffusion-based restoration, these methods often face limitations in generation capability and sampling efficiency. In this work, we present SeedVR, a diffusion transformer designed to handle real-world video restoration with arbitrary length and resolution. The core design of SeedVR lies in the shifted window attention that facilitates effective restoration on long video sequences. SeedVR further supports variable-sized windows near the boundary of both spatial and temporal dimensions, overcoming the resolution constraints of traditional window attention. Equipped with contemporary practices, including causal video autoencoder, mixed image and video training, and progressive training, SeedVR achieves highly-competitive performance on both synthetic and real-world benchmarks, as well as AI-generated videos. Extensive experiments demonstrate SeedVR's superiority over existing methods for generic video restoration.
format Preprint
id arxiv_https___arxiv_org_abs_2501_01320
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration
Wang, Jianyi
Lin, Zhijie
Wei, Meng
Zhao, Yang
Yang, Ceyuan
Xiao, Fei
Loy, Chen Change
Jiang, Lu
Computer Vision and Pattern Recognition
Video restoration poses non-trivial challenges in maintaining fidelity while recovering temporally consistent details from unknown degradations in the wild. Despite recent advances in diffusion-based restoration, these methods often face limitations in generation capability and sampling efficiency. In this work, we present SeedVR, a diffusion transformer designed to handle real-world video restoration with arbitrary length and resolution. The core design of SeedVR lies in the shifted window attention that facilitates effective restoration on long video sequences. SeedVR further supports variable-sized windows near the boundary of both spatial and temporal dimensions, overcoming the resolution constraints of traditional window attention. Equipped with contemporary practices, including causal video autoencoder, mixed image and video training, and progressive training, SeedVR achieves highly-competitive performance on both synthetic and real-world benchmarks, as well as AI-generated videos. Extensive experiments demonstrate SeedVR's superiority over existing methods for generic video restoration.
title SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.01320