Revisiting MLLM Based Image Quality Assessment: Errors and Remedy

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tang, Zhenchen, Yang, Songlin, Peng, Bo, Wang, Zichuan, Dong, Jing
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914150221873152
author Tang, Zhenchen
Yang, Songlin
Peng, Bo
Wang, Zichuan
Dong, Jing
author_facet Tang, Zhenchen
Yang, Songlin
Peng, Bo
Wang, Zichuan
Dong, Jing
contents The rapid progress of multi-modal large language models (MLLMs) has boosted the task of image quality assessment (IQA). However, a key challenge arises from the inherent mismatch between the discrete token outputs of MLLMs and the continuous nature of quality scores required by IQA tasks. This discrepancy significantly hinders the performance of MLLM-based IQA methods. Previous approaches that convert discrete token predictions into continuous scores often suffer from conversion errors. Moreover, the semantic confusion introduced by level tokens (e.g., ``good'') further constrains the performance of MLLMs on IQA tasks and degrades their original capabilities for related tasks. To tackle these problems, we provide a theoretical analysis of the errors inherent in previous approaches and, motivated by this analysis, propose a simple yet effective framework, Q-Scorer. This framework incorporates a lightweight regression module and IQA-specific score tokens into the MLLM pipeline. Extensive experiments demonstrate that Q-Scorer achieves state-of-the-art performance across multiple IQA benchmarks, generalizes well to mixed datasets, and further improves when combined with other methods.
format Preprint
id arxiv_https___arxiv_org_abs_2511_07812
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Revisiting MLLM Based Image Quality Assessment: Errors and Remedy
Tang, Zhenchen
Yang, Songlin
Peng, Bo
Wang, Zichuan
Dong, Jing
Computer Vision and Pattern Recognition
The rapid progress of multi-modal large language models (MLLMs) has boosted the task of image quality assessment (IQA). However, a key challenge arises from the inherent mismatch between the discrete token outputs of MLLMs and the continuous nature of quality scores required by IQA tasks. This discrepancy significantly hinders the performance of MLLM-based IQA methods. Previous approaches that convert discrete token predictions into continuous scores often suffer from conversion errors. Moreover, the semantic confusion introduced by level tokens (e.g., ``good'') further constrains the performance of MLLMs on IQA tasks and degrades their original capabilities for related tasks. To tackle these problems, we provide a theoretical analysis of the errors inherent in previous approaches and, motivated by this analysis, propose a simple yet effective framework, Q-Scorer. This framework incorporates a lightweight regression module and IQA-specific score tokens into the MLLM pipeline. Extensive experiments demonstrate that Q-Scorer achieves state-of-the-art performance across multiple IQA benchmarks, generalizes well to mixed datasets, and further improves when combined with other methods.
title Revisiting MLLM Based Image Quality Assessment: Errors and Remedy
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.07812