Priorformer: A UGC-VQA Method with content and distortion priors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pei, Yajing, Huang, Shiyu, Lu, Yiting, Li, Xin, Chen, Zhibo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909230137606144
author Pei, Yajing
Huang, Shiyu
Lu, Yiting
Li, Xin
Chen, Zhibo
author_facet Pei, Yajing
Huang, Shiyu
Lu, Yiting
Li, Xin
Chen, Zhibo
contents User Generated Content (UGC) videos are susceptible to complicated and variant degradations and contents, which prevents the existing blind video quality assessment (BVQA) models from good performance since the lack of the adapability of distortions and contents. To mitigate this, we propose a novel prior-augmented perceptual vision transformer (PriorFormer) for the BVQA of UGC, which boots its adaptability and representation capability for divergent contents and distortions. Concretely, we introduce two powerful priors, i.e., the content and distortion priors, by extracting the content and distortion embeddings from two pre-trained feature extractors. Then we adopt these two powerful embeddings as the adaptive prior tokens, which are transferred to the vision transformer backbone jointly with implicit quality features. Based on the above strategy, the proposed PriorFormer achieves state-of-the-art performance on three public UGC VQA datasets including KoNViD-1K, LIVE-VQC and YouTube-UGC.
format Preprint
id arxiv_https___arxiv_org_abs_2406_16297
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Priorformer: A UGC-VQA Method with content and distortion priors
Pei, Yajing
Huang, Shiyu
Lu, Yiting
Li, Xin
Chen, Zhibo
Computer Vision and Pattern Recognition
Image and Video Processing
User Generated Content (UGC) videos are susceptible to complicated and variant degradations and contents, which prevents the existing blind video quality assessment (BVQA) models from good performance since the lack of the adapability of distortions and contents. To mitigate this, we propose a novel prior-augmented perceptual vision transformer (PriorFormer) for the BVQA of UGC, which boots its adaptability and representation capability for divergent contents and distortions. Concretely, we introduce two powerful priors, i.e., the content and distortion priors, by extracting the content and distortion embeddings from two pre-trained feature extractors. Then we adopt these two powerful embeddings as the adaptive prior tokens, which are transferred to the vision transformer backbone jointly with implicit quality features. Based on the above strategy, the proposed PriorFormer achieves state-of-the-art performance on three public UGC VQA datasets including KoNViD-1K, LIVE-VQC and YouTube-UGC.
title Priorformer: A UGC-VQA Method with content and distortion priors
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2406.16297