DPC-VQA: Decoupling Quality Perception and Residual Calibration for Video Quality Assessment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Xinyue, Xu, Shubo, Zhang, Zhichao, Cai, Zhaolin, Chen, Yitong, Zhai, Guangtao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914471761412096
author Li, Xinyue
Xu, Shubo
Zhang, Zhichao
Cai, Zhaolin
Chen, Yitong
Zhai, Guangtao
author_facet Li, Xinyue
Xu, Shubo
Zhang, Zhichao
Cai, Zhaolin
Chen, Yitong
Zhai, Guangtao
contents Recent multimodal large language models (MLLMs) have shown promising performance on video quality assessment (VQA) tasks. However, adapting them to new scenarios remains expensive due to large-scale retraining and costly mean opinion score (MOS) annotations. In this paper, we argue that a pretrained MLLM already provides a useful perceptual prior for VQA, and that the main challenge is to efficiently calibrate this prior to the target MOS space. Based on this insight, we propose DPC-VQA, a decoupling perception and calibration framework for video quality assessment. Specifically, DPC-VQA uses a frozen MLLM to provide a base quality estimate and perceptual prior, and employs a lightweight calibration branch to predict a residual correction for target-scenario adaptation. This design avoids costly end-to-end retraining while maintaining reliable performance with lower training and data costs. Extensive experiments on both user-generated content (UGC) and AI-generated content (AIGC) benchmarks show that DPC-VQA achieves competitive performance against representative baselines, while using less than 2% of the trainable parameters of conventional MLLM-based VQA methods and remaining effective with only 20\% of MOS labels. The code will be released upon publication.
format Preprint
id arxiv_https___arxiv_org_abs_2604_12813
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DPC-VQA: Decoupling Quality Perception and Residual Calibration for Video Quality Assessment
Li, Xinyue
Xu, Shubo
Zhang, Zhichao
Cai, Zhaolin
Chen, Yitong
Zhai, Guangtao
Computer Vision and Pattern Recognition
Multimedia
Recent multimodal large language models (MLLMs) have shown promising performance on video quality assessment (VQA) tasks. However, adapting them to new scenarios remains expensive due to large-scale retraining and costly mean opinion score (MOS) annotations. In this paper, we argue that a pretrained MLLM already provides a useful perceptual prior for VQA, and that the main challenge is to efficiently calibrate this prior to the target MOS space. Based on this insight, we propose DPC-VQA, a decoupling perception and calibration framework for video quality assessment. Specifically, DPC-VQA uses a frozen MLLM to provide a base quality estimate and perceptual prior, and employs a lightweight calibration branch to predict a residual correction for target-scenario adaptation. This design avoids costly end-to-end retraining while maintaining reliable performance with lower training and data costs. Extensive experiments on both user-generated content (UGC) and AI-generated content (AIGC) benchmarks show that DPC-VQA achieves competitive performance against representative baselines, while using less than 2% of the trainable parameters of conventional MLLM-based VQA methods and remaining effective with only 20\% of MOS labels. The code will be released upon publication.
title DPC-VQA: Decoupling Quality Perception and Residual Calibration for Video Quality Assessment
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2604.12813