Understanding Degradation with Vision Language Model

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lan, Guanzhou, Liao, Chenyi, Yang, Yuqi, Ma, Qianli, Wang, Zhigang, Wang, Dong, Zhao, Bin, Li, Xuelong
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915774038278144
author Lan, Guanzhou
Liao, Chenyi
Yang, Yuqi
Ma, Qianli
Wang, Zhigang
Wang, Dong
Zhao, Bin
Li, Xuelong
author_facet Lan, Guanzhou
Liao, Chenyi
Yang, Yuqi
Ma, Qianli
Wang, Zhigang
Wang, Dong
Zhao, Bin
Li, Xuelong
contents Understanding visual degradations is a critical yet challenging problem in computer vision. While recent Vision-Language Models (VLMs) excel at qualitative description, they often fall short in understanding the parametric physics underlying image degradations. In this work, we redefine degradation understanding as a hierarchical structured prediction task, necessitating the concurrent estimation of degradation types, parameter keys, and their continuous physical values. Although these sub-tasks operate in disparate spaces, we prove that they can be unified under one autoregressive next-token prediction paradigm, whose error is bounded by the value-space quantization grid. Building on this insight, we introduce DU-VLM, a multimodal chain-of-thought model trained with supervised fine-tuning and reinforcement learning using structured rewards. Furthermore, we show that DU-VLM can serve as a zero-shot controller for pre-trained diffusion models, enabling high-fidelity image restoration without fine-tuning the generative backbone. We also introduce \textbf{DU-110k}, a large-scale dataset comprising 110,000 clean-degraded pairs with grounded physical annotations. Extensive experiments demonstrate that our approach significantly outperforms generalist baselines in both accuracy and robustness, exhibiting generalization to unseen distributions.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04565
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Understanding Degradation with Vision Language Model
Lan, Guanzhou
Liao, Chenyi
Yang, Yuqi
Ma, Qianli
Wang, Zhigang
Wang, Dong
Zhao, Bin
Li, Xuelong
Computer Vision and Pattern Recognition
Understanding visual degradations is a critical yet challenging problem in computer vision. While recent Vision-Language Models (VLMs) excel at qualitative description, they often fall short in understanding the parametric physics underlying image degradations. In this work, we redefine degradation understanding as a hierarchical structured prediction task, necessitating the concurrent estimation of degradation types, parameter keys, and their continuous physical values. Although these sub-tasks operate in disparate spaces, we prove that they can be unified under one autoregressive next-token prediction paradigm, whose error is bounded by the value-space quantization grid. Building on this insight, we introduce DU-VLM, a multimodal chain-of-thought model trained with supervised fine-tuning and reinforcement learning using structured rewards. Furthermore, we show that DU-VLM can serve as a zero-shot controller for pre-trained diffusion models, enabling high-fidelity image restoration without fine-tuning the generative backbone. We also introduce \textbf{DU-110k}, a large-scale dataset comprising 110,000 clean-degraded pairs with grounded physical annotations. Extensive experiments demonstrate that our approach significantly outperforms generalist baselines in both accuracy and robustness, exhibiting generalization to unseen distributions.
title Understanding Degradation with Vision Language Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.04565