UniQA: Unified Vision-Language Pre-training for Image Quality and Aesthetic Assessment

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhou, Hantao, Tang, Longxiang, Yang, Rui, Qin, Guanyi, Zhang, Yan, Li, Yutao, Li, Xiu, Hu, Runze, Zhai, Guangtao
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911053124730880
author Zhou, Hantao
Tang, Longxiang
Yang, Rui
Qin, Guanyi
Zhang, Yan
Li, Yutao
Li, Xiu
Hu, Runze
Zhai, Guangtao
author_facet Zhou, Hantao
Tang, Longxiang
Yang, Rui
Qin, Guanyi
Zhang, Yan
Li, Yutao
Li, Xiu
Hu, Runze
Zhai, Guangtao
contents Image Quality Assessment (IQA) and Image Aesthetic Assessment (IAA) aim to simulate human subjective perception of image visual quality and aesthetic appeal. Despite distinct learning objectives, they have underlying interconnectedness due to consistent human assessment perception. In this paper, we propose Unified vision-language pre-training of Quality and Aesthetics (UniQA}), to extract useful and common representations from two tasks, thereby benefiting them simultaneously. However, the lack of text in the IQA datasets and the textual noise in the IAA datasets pose severe challenges for multimodal pre-training. To address this, we (1) utilize multimodal large language models (MLLMs) to generate high-quality text descriptions; (2) use the generated text for IAA as metadata to purify noisy IAA data. To effectively adapt the pre-trained UniQA to downstream tasks, we further propose a lightweight adapter that utilizes versatile cues to fully exploit the extensive knowledge of the pre-trained model. UniQA demonstrates high competitiveness in various image assessment tasks, including classical IQA and IAA tasks, few-label IQA, and other downstream tasks, showing promise as a foundational assessment model. Codes are available at https://github.com/zht8506/UniQA.
format Preprint
id arxiv_https___arxiv_org_abs_2406_01069
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle UniQA: Unified Vision-Language Pre-training for Image Quality and Aesthetic Assessment
Zhou, Hantao
Tang, Longxiang
Yang, Rui
Qin, Guanyi
Zhang, Yan
Li, Yutao
Li, Xiu
Hu, Runze
Zhai, Guangtao
Computer Vision and Pattern Recognition
Image Quality Assessment (IQA) and Image Aesthetic Assessment (IAA) aim to simulate human subjective perception of image visual quality and aesthetic appeal. Despite distinct learning objectives, they have underlying interconnectedness due to consistent human assessment perception. In this paper, we propose Unified vision-language pre-training of Quality and Aesthetics (UniQA}), to extract useful and common representations from two tasks, thereby benefiting them simultaneously. However, the lack of text in the IQA datasets and the textual noise in the IAA datasets pose severe challenges for multimodal pre-training. To address this, we (1) utilize multimodal large language models (MLLMs) to generate high-quality text descriptions; (2) use the generated text for IAA as metadata to purify noisy IAA data. To effectively adapt the pre-trained UniQA to downstream tasks, we further propose a lightweight adapter that utilizes versatile cues to fully exploit the extensive knowledge of the pre-trained model. UniQA demonstrates high competitiveness in various image assessment tasks, including classical IQA and IAA tasks, few-label IQA, and other downstream tasks, showing promise as a foundational assessment model. Codes are available at https://github.com/zht8506/UniQA.
title UniQA: Unified Vision-Language Pre-training for Image Quality and Aesthetic Assessment
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.01069