Q-Probe: Scaling Image Quality Assessment to High Resolution via Context-Aware Agentic Probing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Xiang, Li, Xueheng, Wang, Yu, He, Xuanhua, Hu, Zhangchi, Yu, Weiwei, Xie, Chengjun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917470520999936
author Li, Xiang
Li, Xueheng
Wang, Yu
He, Xuanhua
Hu, Zhangchi
Yu, Weiwei
Xie, Chengjun
author_facet Li, Xiang
Li, Xueheng
Wang, Yu
He, Xuanhua
Hu, Zhangchi
Yu, Weiwei
Xie, Chengjun
contents Reinforcement Learning (RL) has empowered Multimodal Large Language Models (MLLMs) to achieve superior human preference alignment in Image Quality Assessment (IQA). However, existing RL-based IQA models typically rely on coarse-grained global views, failing to capture subtle local degradations in high-resolution scenarios. While emerging "Thinking with Images" paradigms enable multi-scale visual perception via zoom-in mechanisms, their direct adaptation to IQA induces spurious "cropping-implies-degradation" biases and misinterprets natural depth-of-field as artifacts. To address these challenges, we propose Q-Probe, the first agentic IQA framework designed to scale IQA to high resolution via context-aware probing. First, we construct Vista-Bench, a pioneering benchmark tailored for fine-grained local degradation analysis in high-resolution IQA settings. Furthermore, we propose a three-stage training paradigm that progressively aligns the model with human preferences, while simultaneously eliminating causal bias through a novel context-aware cropping strategy. Extensive experiments demonstrate that Q-Probe achieves state-of-the-art performance in high-resolution settings while maintaining superior efficacy across resolution scales.
format Preprint
id arxiv_https___arxiv_org_abs_2601_15356
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Q-Probe: Scaling Image Quality Assessment to High Resolution via Context-Aware Agentic Probing
Li, Xiang
Li, Xueheng
Wang, Yu
He, Xuanhua
Hu, Zhangchi
Yu, Weiwei
Xie, Chengjun
Image and Video Processing
Artificial Intelligence
Reinforcement Learning (RL) has empowered Multimodal Large Language Models (MLLMs) to achieve superior human preference alignment in Image Quality Assessment (IQA). However, existing RL-based IQA models typically rely on coarse-grained global views, failing to capture subtle local degradations in high-resolution scenarios. While emerging "Thinking with Images" paradigms enable multi-scale visual perception via zoom-in mechanisms, their direct adaptation to IQA induces spurious "cropping-implies-degradation" biases and misinterprets natural depth-of-field as artifacts. To address these challenges, we propose Q-Probe, the first agentic IQA framework designed to scale IQA to high resolution via context-aware probing. First, we construct Vista-Bench, a pioneering benchmark tailored for fine-grained local degradation analysis in high-resolution IQA settings. Furthermore, we propose a three-stage training paradigm that progressively aligns the model with human preferences, while simultaneously eliminating causal bias through a novel context-aware cropping strategy. Extensive experiments demonstrate that Q-Probe achieves state-of-the-art performance in high-resolution settings while maintaining superior efficacy across resolution scales.
title Q-Probe: Scaling Image Quality Assessment to High Resolution via Context-Aware Agentic Probing
topic Image and Video Processing
Artificial Intelligence
url https://arxiv.org/abs/2601.15356