ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liang, Yijun, Li, Ming, Fan, Chenrui, Li, Ziyue, Nguyen, Dang, Cobbina, Kwesi, Bhardwaj, Shweta, Chen, Jiuhai, Liu, Fuxiao, Zhou, Tianyi
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915605252145152
author Liang, Yijun
Li, Ming
Fan, Chenrui
Li, Ziyue
Nguyen, Dang
Cobbina, Kwesi
Bhardwaj, Shweta
Chen, Jiuhai
Liu, Fuxiao
Zhou, Tianyi
author_facet Liang, Yijun
Li, Ming
Fan, Chenrui
Li, Ziyue
Nguyen, Dang
Cobbina, Kwesi
Bhardwaj, Shweta
Chen, Jiuhai
Liu, Fuxiao
Zhou, Tianyi
contents Color plays an important role in human perception and usually provides critical clues in visual reasoning. However, it is unclear whether and how vision-language models (VLMs) can perceive, understand, and leverage color as humans. This paper introduces ColorBench, an innovative benchmark meticulously crafted to assess the capabilities of VLMs in color understanding, including color perception, reasoning, and robustness. By curating a suite of diverse test scenarios, with grounding in real applications, ColorBench evaluates how these models perceive colors, infer meanings from color-based cues, and maintain consistent performance under varying color transformations. Through an extensive evaluation of 32 VLMs with varying language models and vision encoders, our paper reveals some undiscovered findings: (i) The scaling law (larger models are better) still holds on ColorBench, while the language model plays a more important role than the vision encoder. (ii) However, the performance gaps across models are relatively small, indicating that color understanding has been largely neglected by existing VLMs. (iii) CoT reasoning improves color understanding accuracies and robustness, though they are vision-centric tasks. (iv) Color clues are indeed leveraged by VLMs on ColorBench but they can also mislead models in some tasks. These findings highlight the critical limitations of current VLMs and underscore the need to enhance color comprehension. Our ColorBenchcan serve as a foundational tool for advancing the study of human-level color understanding of multimodal AI.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10514
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness
Liang, Yijun
Li, Ming
Fan, Chenrui
Li, Ziyue
Nguyen, Dang
Cobbina, Kwesi
Bhardwaj, Shweta
Chen, Jiuhai
Liu, Fuxiao
Zhou, Tianyi
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Color plays an important role in human perception and usually provides critical clues in visual reasoning. However, it is unclear whether and how vision-language models (VLMs) can perceive, understand, and leverage color as humans. This paper introduces ColorBench, an innovative benchmark meticulously crafted to assess the capabilities of VLMs in color understanding, including color perception, reasoning, and robustness. By curating a suite of diverse test scenarios, with grounding in real applications, ColorBench evaluates how these models perceive colors, infer meanings from color-based cues, and maintain consistent performance under varying color transformations. Through an extensive evaluation of 32 VLMs with varying language models and vision encoders, our paper reveals some undiscovered findings: (i) The scaling law (larger models are better) still holds on ColorBench, while the language model plays a more important role than the vision encoder. (ii) However, the performance gaps across models are relatively small, indicating that color understanding has been largely neglected by existing VLMs. (iii) CoT reasoning improves color understanding accuracies and robustness, though they are vision-centric tasks. (iv) Color clues are indeed leveraged by VLMs on ColorBench but they can also mislead models in some tasks. These findings highlight the critical limitations of current VLMs and underscore the need to enhance color comprehension. Our ColorBenchcan serve as a foundational tool for advancing the study of human-level color understanding of multimodal AI.
title ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2504.10514