Grounding-IQA: Grounding Multimodal Language Model for Image Quality Assessment

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Zheng, Zhang, Xun, Li, Wenbo, Pei, Renjing, Song, Fenglong, Min, Xiongkuo, Liu, Xiaohong, Yuan, Xin, Guo, Yong, Zhang, Yulun
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918361809551360
author Chen, Zheng
Zhang, Xun
Li, Wenbo
Pei, Renjing
Song, Fenglong
Min, Xiongkuo
Liu, Xiaohong
Yuan, Xin
Guo, Yong
Zhang, Yulun
author_facet Chen, Zheng
Zhang, Xun
Li, Wenbo
Pei, Renjing
Song, Fenglong
Min, Xiongkuo
Liu, Xiaohong
Yuan, Xin
Guo, Yong
Zhang, Yulun
contents The development of multimodal large language models (MLLMs) enables the evaluation of image quality through natural language descriptions. This advancement allows for more detailed assessments. However, these MLLM-based IQA methods primarily rely on general contextual descriptions, sometimes limiting fine-grained quality assessment. To address this limitation, we introduce a new image quality assessment (IQA) task paradigm, **grounding-IQA**. This paradigm integrates multimodal referring and grounding with IQA to realize more fine-grained quality perception, thereby extending existing IQA. Specifically, grounding-IQA comprises two subtasks: grounding-IQA-description (GIQA-DES) and visual question answering (GIQA-VQA). GIQA-DES involves detailed descriptions with precise locations (e.g., bounding boxes), while GIQA-VQA focuses on quality QA for local regions. To realize grounding-IQA, we construct a corresponding dataset, GIQA-160K, through our proposed automated annotation pipeline. Furthermore, we develop a well-designed benchmark, GIQA-Bench. The benchmark evaluates the grounding-IQA performance from three perspectives: description quality, VQA accuracy, and grounding precision. Experiments demonstrate that our proposed method facilitates the more fine-grained IQA application. Code: https://github.com/zhengchen1999/Grounding-IQA.
format Preprint
id arxiv_https___arxiv_org_abs_2411_17237
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Grounding-IQA: Grounding Multimodal Language Model for Image Quality Assessment
Chen, Zheng
Zhang, Xun
Li, Wenbo
Pei, Renjing
Song, Fenglong
Min, Xiongkuo
Liu, Xiaohong
Yuan, Xin
Guo, Yong
Zhang, Yulun
Computer Vision and Pattern Recognition
The development of multimodal large language models (MLLMs) enables the evaluation of image quality through natural language descriptions. This advancement allows for more detailed assessments. However, these MLLM-based IQA methods primarily rely on general contextual descriptions, sometimes limiting fine-grained quality assessment. To address this limitation, we introduce a new image quality assessment (IQA) task paradigm, **grounding-IQA**. This paradigm integrates multimodal referring and grounding with IQA to realize more fine-grained quality perception, thereby extending existing IQA. Specifically, grounding-IQA comprises two subtasks: grounding-IQA-description (GIQA-DES) and visual question answering (GIQA-VQA). GIQA-DES involves detailed descriptions with precise locations (e.g., bounding boxes), while GIQA-VQA focuses on quality QA for local regions. To realize grounding-IQA, we construct a corresponding dataset, GIQA-160K, through our proposed automated annotation pipeline. Furthermore, we develop a well-designed benchmark, GIQA-Bench. The benchmark evaluates the grounding-IQA performance from three perspectives: description quality, VQA accuracy, and grounding precision. Experiments demonstrate that our proposed method facilitates the more fine-grained IQA application. Code: https://github.com/zhengchen1999/Grounding-IQA.
title Grounding-IQA: Grounding Multimodal Language Model for Image Quality Assessment
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.17237