Q-Bench-Portrait: Benchmarking Multimodal Large Language Models on Portrait Image Quality Perception
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914280660533248 |
|---|---|
| author | Wu, Sijing Li, Yunhao Zhang, Zicheng Jia, Qi Li, Xinyue Duan, Huiyu Min, Xiongkuo Zhai, Guangtao |
| author_facet | Wu, Sijing Li, Yunhao Zhang, Zicheng Jia, Qi Li, Xinyue Duan, Huiyu Min, Xiongkuo Zhai, Guangtao |
| contents | Recent advances in multimodal large language models (MLLMs) have demonstrated impressive performance on existing low-level vision benchmarks, which primarily focus on generic images. However, their capabilities to perceive and assess portrait images, a domain characterized by distinct structural and perceptual properties, remain largely underexplored. To this end, we introduce Q-Bench-Portrait, the first holistic benchmark specifically designed for portrait image quality perception, comprising 2,765 image-question-answer triplets and featuring (1) diverse portrait image sources, including natural, synthetic distortion, AI-generated, artistic, and computer graphics images; (2) comprehensive quality dimensions, covering technical distortions, AIGC-specific distortions, and aesthetics; and (3) a range of question formats, including single-choice, multiple-choice, true/false, and open-ended questions, at both global and local levels. Based on Q-Bench-Portrait, we evaluate 20 open-source and 5 closed-source MLLMs, revealing that although current models demonstrate some competence in portrait image perception, their performance remains limited and imprecise, with a clear gap relative to human judgments. We hope that the proposed benchmark will foster further research into enhancing the portrait image perception capabilities of both general-purpose and domain-specific MLLMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_18346 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Q-Bench-Portrait: Benchmarking Multimodal Large Language Models on Portrait Image Quality Perception Wu, Sijing Li, Yunhao Zhang, Zicheng Jia, Qi Li, Xinyue Duan, Huiyu Min, Xiongkuo Zhai, Guangtao Computer Vision and Pattern Recognition Recent advances in multimodal large language models (MLLMs) have demonstrated impressive performance on existing low-level vision benchmarks, which primarily focus on generic images. However, their capabilities to perceive and assess portrait images, a domain characterized by distinct structural and perceptual properties, remain largely underexplored. To this end, we introduce Q-Bench-Portrait, the first holistic benchmark specifically designed for portrait image quality perception, comprising 2,765 image-question-answer triplets and featuring (1) diverse portrait image sources, including natural, synthetic distortion, AI-generated, artistic, and computer graphics images; (2) comprehensive quality dimensions, covering technical distortions, AIGC-specific distortions, and aesthetics; and (3) a range of question formats, including single-choice, multiple-choice, true/false, and open-ended questions, at both global and local levels. Based on Q-Bench-Portrait, we evaluate 20 open-source and 5 closed-source MLLMs, revealing that although current models demonstrate some competence in portrait image perception, their performance remains limited and imprecise, with a clear gap relative to human judgments. We hope that the proposed benchmark will foster further research into enhancing the portrait image perception capabilities of both general-purpose and domain-specific MLLMs. |
| title | Q-Bench-Portrait: Benchmarking Multimodal Large Language Models on Portrait Image Quality Perception |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2601.18346 |