iDETEX: Empowering MLLMs for Intelligent DETailed EXplainable IQA

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Zhaoran, Yue, Xinli, Sun, Jianhui, Xie, Yuhao, Shao, Tao, Yao, Liangchao, Xia, Fan, Deng, Yuetang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917027369713664
author Zhao, Zhaoran
Yue, Xinli
Sun, Jianhui
Xie, Yuhao
Shao, Tao
Yao, Liangchao
Xia, Fan
Deng, Yuetang
author_facet Zhao, Zhaoran
Yue, Xinli
Sun, Jianhui
Xie, Yuhao
Shao, Tao
Yao, Liangchao
Xia, Fan
Deng, Yuetang
contents Image Quality Assessment (IQA) has progressed from scalar quality prediction to more interpretable, human-aligned evaluation paradigms. In this work, we address the emerging challenge of detailed and explainable IQA by proposing iDETEX-a unified multimodal large language model (MLLM) capable of simultaneously performing three key tasks: quality grounding, perception, and description. To facilitate efficient and generalizable training across these heterogeneous subtasks, we design a suite of task-specific offline augmentation modules and a data mixing strategy. These are further complemented by online enhancement strategies to fully exploit multi-sourced supervision. We validate our approach on the large-scale ViDA-UGC benchmark, where iDETEX achieves state-of-the-art performance across all subtasks. Our model ranks first in the ICCV MIPI 2025 Detailed Image Quality Assessment Challenge, demonstrating its effectiveness and robustness in delivering accurate and interpretable quality assessments.
format Preprint
id arxiv_https___arxiv_org_abs_2510_17332
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle iDETEX: Empowering MLLMs for Intelligent DETailed EXplainable IQA
Zhao, Zhaoran
Yue, Xinli
Sun, Jianhui
Xie, Yuhao
Shao, Tao
Yao, Liangchao
Xia, Fan
Deng, Yuetang
Computer Vision and Pattern Recognition
Image Quality Assessment (IQA) has progressed from scalar quality prediction to more interpretable, human-aligned evaluation paradigms. In this work, we address the emerging challenge of detailed and explainable IQA by proposing iDETEX-a unified multimodal large language model (MLLM) capable of simultaneously performing three key tasks: quality grounding, perception, and description. To facilitate efficient and generalizable training across these heterogeneous subtasks, we design a suite of task-specific offline augmentation modules and a data mixing strategy. These are further complemented by online enhancement strategies to fully exploit multi-sourced supervision. We validate our approach on the large-scale ViDA-UGC benchmark, where iDETEX achieves state-of-the-art performance across all subtasks. Our model ranks first in the ICCV MIPI 2025 Detailed Image Quality Assessment Challenge, demonstrating its effectiveness and robustness in delivering accurate and interpretable quality assessments.
title iDETEX: Empowering MLLMs for Intelligent DETailed EXplainable IQA
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.17332