HanMoVLM: Large Vision-Language Models for Professional Artistic Painting Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Hongji, Zhou, Yucheng, Han, Wencheng, Li, Songlian, Zhao, Xiaotong, Shen, Jianbing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908879166636032
author Yang, Hongji
Zhou, Yucheng
Han, Wencheng
Li, Songlian
Zhao, Xiaotong
Shen, Jianbing
author_facet Yang, Hongji
Zhou, Yucheng
Han, Wencheng
Li, Songlian
Zhao, Xiaotong
Shen, Jianbing
contents While Large Vision-Language Models (VLMs) demonstrate impressive general visual capabilities, they remain artistically blind and unable to offer professional evaluation of artworks within specific artistic domains like human experts. To bridge this gap, we transform VLMs into experts capable of professional-grade painting evaluation in the Chinese Artistic Domain, which is more abstract and demands extensive artistic training for evaluation. We introduce HanMo-Bench, a new dataset that features authentic auction-grade masterpieces and AI-generated works, grounded in real-world market valuations. To realize the rigorous judgment, we propose the HanMoVLM and construct a Chain-of-Thought (CoT) validated by experts. This CoT guides the model to perform expert-level reasoning: from content identification and Region of Interest (RoI) localization to professional evaluation, guided by both theme-specific evaluation and typical three-tier evaluation in Chinese paintings. Furthermore, we design a reward function to refine the reasoning process of the HanMoVLM to improve the accuracy. We demonstrate that HanMoVLM can serve as a critical backbone for Test-time Scaling in image generation. By acting as a high-quality verifier, HanMoVLM enables generative models to select the most artistically superior outputs from multiple candidates. Experimental results and human studies confirm that the proposed HanMoVLM effectively bridges the gap, achieving a high consistency with professional experts and significantly improving the quality of Chinese Painting generation.
format Preprint
id arxiv_https___arxiv_org_abs_2603_10814
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HanMoVLM: Large Vision-Language Models for Professional Artistic Painting Evaluation
Yang, Hongji
Zhou, Yucheng
Han, Wencheng
Li, Songlian
Zhao, Xiaotong
Shen, Jianbing
Computer Vision and Pattern Recognition
While Large Vision-Language Models (VLMs) demonstrate impressive general visual capabilities, they remain artistically blind and unable to offer professional evaluation of artworks within specific artistic domains like human experts. To bridge this gap, we transform VLMs into experts capable of professional-grade painting evaluation in the Chinese Artistic Domain, which is more abstract and demands extensive artistic training for evaluation. We introduce HanMo-Bench, a new dataset that features authentic auction-grade masterpieces and AI-generated works, grounded in real-world market valuations. To realize the rigorous judgment, we propose the HanMoVLM and construct a Chain-of-Thought (CoT) validated by experts. This CoT guides the model to perform expert-level reasoning: from content identification and Region of Interest (RoI) localization to professional evaluation, guided by both theme-specific evaluation and typical three-tier evaluation in Chinese paintings. Furthermore, we design a reward function to refine the reasoning process of the HanMoVLM to improve the accuracy. We demonstrate that HanMoVLM can serve as a critical backbone for Test-time Scaling in image generation. By acting as a high-quality verifier, HanMoVLM enables generative models to select the most artistically superior outputs from multiple candidates. Experimental results and human studies confirm that the proposed HanMoVLM effectively bridges the gap, achieving a high consistency with professional experts and significantly improving the quality of Chinese Painting generation.
title HanMoVLM: Large Vision-Language Models for Professional Artistic Painting Evaluation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.10814