V-Bridge: Bridging Video Generative Priors to Versatile Few-shot Image Restoration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Shenghe, Jiang, Junpeng, Li, Wenbo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914391305224192
author Zheng, Shenghe
Jiang, Junpeng
Li, Wenbo
author_facet Zheng, Shenghe
Jiang, Junpeng
Li, Wenbo
contents Large-scale video generative models are trained on vast and diverse visual data, enabling them to internalize rich structural, semantic, and dynamic priors of the visual world. While these models have demonstrated impressive generative capability, their potential as general-purpose visual learners remains largely untapped. In this work, we introduce V-Bridge, a framework that bridges this latent capacity to versatile few-shot image restoration tasks. We reinterpret image restoration not as a static regression problem, but as a progressive generative process, and leverage video models to simulate the gradual refinement from degraded inputs to high-fidelity outputs. Surprisingly, with only 1,000 multi-task training samples (less than 2% of existing restoration methods), pretrained video models can be induced to perform competitive image restoration, achieving multiple tasks with a single model, rivaling specialized architectures designed explicitly for this purpose. Our findings reveal that video generative models implicitly learn powerful and transferable restoration priors that can be activated with only extremely limited data, challenging the traditional boundary between generative modeling and low-level vision, and opening a new design paradigm for foundation models in visual tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2603_13089
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle V-Bridge: Bridging Video Generative Priors to Versatile Few-shot Image Restoration
Zheng, Shenghe
Jiang, Junpeng
Li, Wenbo
Computer Vision and Pattern Recognition
Large-scale video generative models are trained on vast and diverse visual data, enabling them to internalize rich structural, semantic, and dynamic priors of the visual world. While these models have demonstrated impressive generative capability, their potential as general-purpose visual learners remains largely untapped. In this work, we introduce V-Bridge, a framework that bridges this latent capacity to versatile few-shot image restoration tasks. We reinterpret image restoration not as a static regression problem, but as a progressive generative process, and leverage video models to simulate the gradual refinement from degraded inputs to high-fidelity outputs. Surprisingly, with only 1,000 multi-task training samples (less than 2% of existing restoration methods), pretrained video models can be induced to perform competitive image restoration, achieving multiple tasks with a single model, rivaling specialized architectures designed explicitly for this purpose. Our findings reveal that video generative models implicitly learn powerful and transferable restoration priors that can be activated with only extremely limited data, challenging the traditional boundary between generative modeling and low-level vision, and opening a new design paradigm for foundation models in visual tasks.
title V-Bridge: Bridging Video Generative Priors to Versatile Few-shot Image Restoration
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.13089