Enhancing Image Quality Assessment Ability of LMMs via Retrieval-Augmented Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Kang, Duan, Huiyu, Zhang, Zicheng, Zhu, Yucheng, Zhao, Jun, Min, Xiongkuo, Wang, Jia, Zhai, Guangtao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908761456640000
author Fu, Kang
Duan, Huiyu
Zhang, Zicheng
Zhu, Yucheng
Zhao, Jun
Min, Xiongkuo
Wang, Jia
Zhai, Guangtao
author_facet Fu, Kang
Duan, Huiyu
Zhang, Zicheng
Zhu, Yucheng
Zhao, Jun
Min, Xiongkuo
Wang, Jia
Zhai, Guangtao
contents Large Multimodal Models (LMMs) have recently shown remarkable promise in low-level visual perception tasks, particularly in Image Quality Assessment (IQA), demonstrating strong zero-shot capability. However, achieving state-of-the-art performance often requires computationally expensive fine-tuning methods, which aim to align the distribution of quality-related token in output with image quality levels. Inspired by recent training-free works for LMM, we introduce IQARAG, a novel, training-free framework that enhances LMMs' IQA ability. IQARAG leverages Retrieval-Augmented Generation (RAG) to retrieve some semantically similar but quality-variant reference images with corresponding Mean Opinion Scores (MOSs) for input image. These retrieved images and input image are integrated into a specific prompt. Retrieved images provide the LMM with a visual perception anchor for IQA task. IQARAG contains three key phases: Retrieval Feature Extraction, Image Retrieval, and Integration & Quality Score Generation. Extensive experiments across multiple diverse IQA datasets, including KADID, KonIQ, LIVE Challenge, and SPAQ, demonstrate that the proposed IQARAG effectively boosts the IQA performance of LMMs, offering a resource-efficient alternative to fine-tuning for quality assessment.
format Preprint
id arxiv_https___arxiv_org_abs_2601_08311
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Enhancing Image Quality Assessment Ability of LMMs via Retrieval-Augmented Generation
Fu, Kang
Duan, Huiyu
Zhang, Zicheng
Zhu, Yucheng
Zhao, Jun
Min, Xiongkuo
Wang, Jia
Zhai, Guangtao
Computer Vision and Pattern Recognition
Artificial Intelligence
Large Multimodal Models (LMMs) have recently shown remarkable promise in low-level visual perception tasks, particularly in Image Quality Assessment (IQA), demonstrating strong zero-shot capability. However, achieving state-of-the-art performance often requires computationally expensive fine-tuning methods, which aim to align the distribution of quality-related token in output with image quality levels. Inspired by recent training-free works for LMM, we introduce IQARAG, a novel, training-free framework that enhances LMMs' IQA ability. IQARAG leverages Retrieval-Augmented Generation (RAG) to retrieve some semantically similar but quality-variant reference images with corresponding Mean Opinion Scores (MOSs) for input image. These retrieved images and input image are integrated into a specific prompt. Retrieved images provide the LMM with a visual perception anchor for IQA task. IQARAG contains three key phases: Retrieval Feature Extraction, Image Retrieval, and Integration & Quality Score Generation. Extensive experiments across multiple diverse IQA datasets, including KADID, KonIQ, LIVE Challenge, and SPAQ, demonstrate that the proposed IQARAG effectively boosts the IQA performance of LMMs, offering a resource-efficient alternative to fine-tuning for quality assessment.
title Enhancing Image Quality Assessment Ability of LMMs via Retrieval-Augmented Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2601.08311