Dual-Branch Network for Portrait Image Quality Assessment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Wei, Zhang, Weixia, Jiang, Yanwei, Wu, Haoning, Zhang, Zicheng, Jia, Jun, Zhou, Yingjie, Ji, Zhongpeng, Min, Xiongkuo, Lin, Weisi, Zhai, Guangtao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911876702535680
author Sun, Wei
Zhang, Weixia
Jiang, Yanwei
Wu, Haoning
Zhang, Zicheng
Jia, Jun
Zhou, Yingjie
Ji, Zhongpeng
Min, Xiongkuo
Lin, Weisi
Zhai, Guangtao
author_facet Sun, Wei
Zhang, Weixia
Jiang, Yanwei
Wu, Haoning
Zhang, Zicheng
Jia, Jun
Zhou, Yingjie
Ji, Zhongpeng
Min, Xiongkuo
Lin, Weisi
Zhai, Guangtao
contents Portrait images typically consist of a salient person against diverse backgrounds. With the development of mobile devices and image processing techniques, users can conveniently capture portrait images anytime and anywhere. However, the quality of these portraits may suffer from the degradation caused by unfavorable environmental conditions, subpar photography techniques, and inferior capturing devices. In this paper, we introduce a dual-branch network for portrait image quality assessment (PIQA), which can effectively address how the salient person and the background of a portrait image influence its visual quality. Specifically, we utilize two backbone networks (\textit{i.e.,} Swin Transformer-B) to extract the quality-aware features from the entire portrait image and the facial image cropped from it. To enhance the quality-aware feature representation of the backbones, we pre-train them on the large-scale video quality assessment dataset LSVQ and the large-scale facial image quality assessment dataset GFIQA. Additionally, we leverage LIQE, an image scene classification and quality assessment model, to capture the quality-aware and scene-specific features as the auxiliary features. Finally, we concatenate these features and regress them into quality scores via a multi-perception layer (MLP). We employ the fidelity loss to train the model via a learning-to-rank manner to mitigate inconsistencies in quality scores in the portrait image quality assessment dataset PIQ. Experimental results demonstrate that the proposed model achieves superior performance in the PIQ dataset, validating its effectiveness. The code is available at \url{https://github.com/sunwei925/DN-PIQA.git}.
format Preprint
id arxiv_https___arxiv_org_abs_2405_08555
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dual-Branch Network for Portrait Image Quality Assessment
Sun, Wei
Zhang, Weixia
Jiang, Yanwei
Wu, Haoning
Zhang, Zicheng
Jia, Jun
Zhou, Yingjie
Ji, Zhongpeng
Min, Xiongkuo
Lin, Weisi
Zhai, Guangtao
Computer Vision and Pattern Recognition
Multimedia
Portrait images typically consist of a salient person against diverse backgrounds. With the development of mobile devices and image processing techniques, users can conveniently capture portrait images anytime and anywhere. However, the quality of these portraits may suffer from the degradation caused by unfavorable environmental conditions, subpar photography techniques, and inferior capturing devices. In this paper, we introduce a dual-branch network for portrait image quality assessment (PIQA), which can effectively address how the salient person and the background of a portrait image influence its visual quality. Specifically, we utilize two backbone networks (\textit{i.e.,} Swin Transformer-B) to extract the quality-aware features from the entire portrait image and the facial image cropped from it. To enhance the quality-aware feature representation of the backbones, we pre-train them on the large-scale video quality assessment dataset LSVQ and the large-scale facial image quality assessment dataset GFIQA. Additionally, we leverage LIQE, an image scene classification and quality assessment model, to capture the quality-aware and scene-specific features as the auxiliary features. Finally, we concatenate these features and regress them into quality scores via a multi-perception layer (MLP). We employ the fidelity loss to train the model via a learning-to-rank manner to mitigate inconsistencies in quality scores in the portrait image quality assessment dataset PIQ. Experimental results demonstrate that the proposed model achieves superior performance in the PIQ dataset, validating its effectiveness. The code is available at \url{https://github.com/sunwei925/DN-PIQA.git}.
title Dual-Branch Network for Portrait Image Quality Assessment
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2405.08555