ViDA-UGC: Detailed Image Quality Analysis via Visual Distortion Assessment for UGC Images

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liao, Wenjie, Yuan, Jieyu, Xu, Yifang, Guo, Chunle, Zhang, Zilong, Li, Jihong, Fu, Jiachen, Fan, Haotian, Li, Tao, Cui, Junhui, Li, Chongyi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916905083731968
author Liao, Wenjie
Yuan, Jieyu
Xu, Yifang
Guo, Chunle
Zhang, Zilong
Li, Jihong
Fu, Jiachen
Fan, Haotian
Li, Tao
Cui, Junhui
Li, Chongyi
author_facet Liao, Wenjie
Yuan, Jieyu
Xu, Yifang
Guo, Chunle
Zhang, Zilong
Li, Jihong
Fu, Jiachen
Fan, Haotian
Li, Tao
Cui, Junhui
Li, Chongyi
contents Recent advances in Multimodal Large Language Models (MLLMs) have introduced a paradigm shift for Image Quality Assessment (IQA) from unexplainable image quality scoring to explainable IQA, demonstrating practical applications like quality control and optimization guidance. However, current explainable IQA methods not only inadequately use the same distortion criteria to evaluate both User-Generated Content (UGC) and AI-Generated Content (AIGC) images, but also lack detailed quality analysis for monitoring image quality and guiding image restoration. In this study, we establish the first large-scale Visual Distortion Assessment Instruction Tuning Dataset for UGC images, termed ViDA-UGC, which comprises 11K images with fine-grained quality grounding, detailed quality perception, and reasoning quality description data. This dataset is constructed through a distortion-oriented pipeline, which involves human subject annotation and a Chain-of-Thought (CoT) assessment framework. This framework guides GPT-4o to generate quality descriptions by identifying and analyzing UGC distortions, which helps capturing rich low-level visual features that inherently correlate with distortion patterns. Moreover, we carefully select 476 images with corresponding 6,149 question answer pairs from ViDA-UGC and invite a professional team to ensure the accuracy and quality of GPT-generated information. The selected and revised data further contribute to the first UGC distortion assessment benchmark, termed ViDA-UGC-Bench. Experimental results demonstrate the effectiveness of the ViDA-UGC and CoT framework for consistently enhancing various image quality analysis abilities across multiple base MLLMs on ViDA-UGC-Bench and Q-Bench, even surpassing GPT-4o.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12605
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ViDA-UGC: Detailed Image Quality Analysis via Visual Distortion Assessment for UGC Images
Liao, Wenjie
Yuan, Jieyu
Xu, Yifang
Guo, Chunle
Zhang, Zilong
Li, Jihong
Fu, Jiachen
Fan, Haotian
Li, Tao
Cui, Junhui
Li, Chongyi
Computer Vision and Pattern Recognition
Recent advances in Multimodal Large Language Models (MLLMs) have introduced a paradigm shift for Image Quality Assessment (IQA) from unexplainable image quality scoring to explainable IQA, demonstrating practical applications like quality control and optimization guidance. However, current explainable IQA methods not only inadequately use the same distortion criteria to evaluate both User-Generated Content (UGC) and AI-Generated Content (AIGC) images, but also lack detailed quality analysis for monitoring image quality and guiding image restoration. In this study, we establish the first large-scale Visual Distortion Assessment Instruction Tuning Dataset for UGC images, termed ViDA-UGC, which comprises 11K images with fine-grained quality grounding, detailed quality perception, and reasoning quality description data. This dataset is constructed through a distortion-oriented pipeline, which involves human subject annotation and a Chain-of-Thought (CoT) assessment framework. This framework guides GPT-4o to generate quality descriptions by identifying and analyzing UGC distortions, which helps capturing rich low-level visual features that inherently correlate with distortion patterns. Moreover, we carefully select 476 images with corresponding 6,149 question answer pairs from ViDA-UGC and invite a professional team to ensure the accuracy and quality of GPT-generated information. The selected and revised data further contribute to the first UGC distortion assessment benchmark, termed ViDA-UGC-Bench. Experimental results demonstrate the effectiveness of the ViDA-UGC and CoT framework for consistently enhancing various image quality analysis abilities across multiple base MLLMs on ViDA-UGC-Bench and Q-Bench, even surpassing GPT-4o.
title ViDA-UGC: Detailed Image Quality Analysis via Visual Distortion Assessment for UGC Images
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.12605