Color Names in Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gomez-Villa, Alexandra, Hernández-Cámara, Pablo, Butt, Muhammad Atif, Laparra, Valero, Malo, Jesus, Vazquez-Corral, Javier
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912608778452992
author Gomez-Villa, Alexandra
Hernández-Cámara, Pablo
Butt, Muhammad Atif
Laparra, Valero
Malo, Jesus
Vazquez-Corral, Javier
author_facet Gomez-Villa, Alexandra
Hernández-Cámara, Pablo
Butt, Muhammad Atif
Laparra, Valero
Malo, Jesus
Vazquez-Corral, Javier
contents Color serves as a fundamental dimension of human visual perception and a primary means of communicating about objects and scenes. As vision-language models (VLMs) become increasingly prevalent, understanding whether they name colors like humans is crucial for effective human-AI interaction. We present the first systematic evaluation of color naming capabilities across VLMs, replicating classic color naming methodologies using 957 color samples across five representative models. Our results show that while VLMs achieve high accuracy on prototypical colors from classical studies, performance drops significantly on expanded, non-prototypical color sets. We identify 21 common color terms that consistently emerge across all models, revealing two distinct approaches: constrained models using predominantly basic terms versus expansive models employing systematic lightness modifiers. Cross-linguistic analysis across nine languages demonstrates severe training imbalances favoring English and Chinese, with hue serving as the primary driver of color naming decisions. Finally, ablation studies reveal that language model architecture significantly influences color naming independent of visual processing capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22524
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Color Names in Vision-Language Models
Gomez-Villa, Alexandra
Hernández-Cámara, Pablo
Butt, Muhammad Atif
Laparra, Valero
Malo, Jesus
Vazquez-Corral, Javier
Computer Vision and Pattern Recognition
Color serves as a fundamental dimension of human visual perception and a primary means of communicating about objects and scenes. As vision-language models (VLMs) become increasingly prevalent, understanding whether they name colors like humans is crucial for effective human-AI interaction. We present the first systematic evaluation of color naming capabilities across VLMs, replicating classic color naming methodologies using 957 color samples across five representative models. Our results show that while VLMs achieve high accuracy on prototypical colors from classical studies, performance drops significantly on expanded, non-prototypical color sets. We identify 21 common color terms that consistently emerge across all models, revealing two distinct approaches: constrained models using predominantly basic terms versus expansive models employing systematic lightness modifiers. Cross-linguistic analysis across nine languages demonstrates severe training imbalances favoring English and Chinese, with hue serving as the primary driver of color naming decisions. Finally, ablation studies reveal that language model architecture significantly influences color naming independent of visual processing capabilities.
title Color Names in Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.22524