Context and Pixel Aware Large Language Model for Video Quality Assessment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wen, Wen, Wu, Yaohong, Sheng, Yue, Birkbeck, Neil, Adsumilli, Balu, Wang, Yilin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910202225229824
author Wen, Wen
Wu, Yaohong
Sheng, Yue
Birkbeck, Neil
Adsumilli, Balu
Wang, Yilin
author_facet Wen, Wen
Wu, Yaohong
Sheng, Yue
Birkbeck, Neil
Adsumilli, Balu
Wang, Yilin
contents Video quality assessment (VQA) is a challenging research topic with broad applications. Traditional hand-crafted and discriminative learning-based VQA models mainly focus on pixel-level distortions and lack contextual understanding, while recent multimodal large language models (MLLMs) struggle with sensitivity to small distortions or handle quality scoring and description as separate tasks. To address these shortcomings, we introduce CP-LLM: a Context- and Pixel-aware Large Language Model. CP-LLM is a novel multimodal LLM architecture featuring dual vision encoders designed to independently analyze perceptual quality at both high-level (video context) and low-level (pixel distortion) granularity, along with a language decoder that subsequently reasons about the interplay between these aspects. This design enables CP-LLM to simultaneously produce robust quality scores and interpretable quality descriptions, with enhanced sensitivity to pixel distortions (e.g., compression artifacts). Experiment results demonstrate that CP-LLM achieves state-of-the-art cross-dataset performance on VQA benchmarks and superior robustness to pixel distortions.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16025
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Context and Pixel Aware Large Language Model for Video Quality Assessment
Wen, Wen
Wu, Yaohong
Sheng, Yue
Birkbeck, Neil
Adsumilli, Balu
Wang, Yilin
Computer Vision and Pattern Recognition
Multimedia
Image and Video Processing
Video quality assessment (VQA) is a challenging research topic with broad applications. Traditional hand-crafted and discriminative learning-based VQA models mainly focus on pixel-level distortions and lack contextual understanding, while recent multimodal large language models (MLLMs) struggle with sensitivity to small distortions or handle quality scoring and description as separate tasks. To address these shortcomings, we introduce CP-LLM: a Context- and Pixel-aware Large Language Model. CP-LLM is a novel multimodal LLM architecture featuring dual vision encoders designed to independently analyze perceptual quality at both high-level (video context) and low-level (pixel distortion) granularity, along with a language decoder that subsequently reasons about the interplay between these aspects. This design enables CP-LLM to simultaneously produce robust quality scores and interpretable quality descriptions, with enhanced sensitivity to pixel distortions (e.g., compression artifacts). Experiment results demonstrate that CP-LLM achieves state-of-the-art cross-dataset performance on VQA benchmarks and superior robustness to pixel distortions.
title Context and Pixel Aware Large Language Model for Video Quality Assessment
topic Computer Vision and Pattern Recognition
Multimedia
Image and Video Processing
url https://arxiv.org/abs/2505.16025