A Comprehensive Study of Multimodal Large Language Models for Image Quality Assessment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Tianhe, Ma, Kede, Liang, Jie, Yang, Yujiu, Zhang, Lei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909250085715968
author Wu, Tianhe
Ma, Kede
Liang, Jie
Yang, Yujiu
Zhang, Lei
author_facet Wu, Tianhe
Ma, Kede
Liang, Jie
Yang, Yujiu
Zhang, Lei
contents While Multimodal Large Language Models (MLLMs) have experienced significant advancement in visual understanding and reasoning, their potential to serve as powerful, flexible, interpretable, and text-driven models for Image Quality Assessment (IQA) remains largely unexplored. In this paper, we conduct a comprehensive and systematic study of prompting MLLMs for IQA. We first investigate nine prompting systems for MLLMs as the combinations of three standardized testing procedures in psychophysics (i.e., the single-stimulus, double-stimulus, and multiple-stimulus methods) and three popular prompting strategies in natural language processing (i.e., the standard, in-context, and chain-of-thought prompting). We then present a difficult sample selection procedure, taking into account sample diversity and uncertainty, to further challenge MLLMs equipped with the respective optimal prompting systems. We assess three open-source and one closed-source MLLMs on several visual attributes of image quality (e.g., structural and textural distortions, geometric transformations, and color differences) in both full-reference and no-reference scenarios. Experimental results show that only the closed-source GPT-4V provides a reasonable account for human perception of image quality, but is weak at discriminating fine-grained quality variations (e.g., color differences) and at comparing visual quality of multiple images, tasks humans can perform effortlessly.
format Preprint
id arxiv_https___arxiv_org_abs_2403_10854
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Comprehensive Study of Multimodal Large Language Models for Image Quality Assessment
Wu, Tianhe
Ma, Kede
Liang, Jie
Yang, Yujiu
Zhang, Lei
Computer Vision and Pattern Recognition
While Multimodal Large Language Models (MLLMs) have experienced significant advancement in visual understanding and reasoning, their potential to serve as powerful, flexible, interpretable, and text-driven models for Image Quality Assessment (IQA) remains largely unexplored. In this paper, we conduct a comprehensive and systematic study of prompting MLLMs for IQA. We first investigate nine prompting systems for MLLMs as the combinations of three standardized testing procedures in psychophysics (i.e., the single-stimulus, double-stimulus, and multiple-stimulus methods) and three popular prompting strategies in natural language processing (i.e., the standard, in-context, and chain-of-thought prompting). We then present a difficult sample selection procedure, taking into account sample diversity and uncertainty, to further challenge MLLMs equipped with the respective optimal prompting systems. We assess three open-source and one closed-source MLLMs on several visual attributes of image quality (e.g., structural and textural distortions, geometric transformations, and color differences) in both full-reference and no-reference scenarios. Experimental results show that only the closed-source GPT-4V provides a reasonable account for human perception of image quality, but is weak at discriminating fine-grained quality variations (e.g., color differences) and at comparing visual quality of multiple images, tasks humans can perform effortlessly.
title A Comprehensive Study of Multimodal Large Language Models for Image Quality Assessment
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.10854