When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Nemitz, Jonathan, Eickhoff, Carsten, Li, Junyi Jessy, Mahowald, Kyle, Golovanevsky, Michal, Rudman, William
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908945677811712
author Nemitz, Jonathan
Eickhoff, Carsten
Li, Junyi Jessy
Mahowald, Kyle
Golovanevsky, Michal
Rudman, William
author_facet Nemitz, Jonathan
Eickhoff, Carsten
Li, Junyi Jessy
Mahowald, Kyle
Golovanevsky, Michal
Rudman, William
contents Understanding when Vision-Language Models (VLMs) will behave unexpectedly, whether models can reliably predict their own behavior, and if models adhere to their introspective reasoning are central challenges for trustworthy deployment. To study this, we introduce the Graded Color Attribution (GCA) dataset, a controlled benchmark designed to elicit decision rules and evaluate participant faithfulness to these rules. GCA consists of line drawings that vary pixel-level color coverage across three conditions: world-knowledge recolorings, counterfactual recolorings, and shapes with no color priors. Using GCA, both VLMs and human participants establish a threshold: the minimum percentage of pixels of a given color an object must have to receive that color label. We then compare these rules with their subsequent color attribution decisions. Our findings reveal that models systematically violate their own introspective rules. For example, GPT-5-mini violates its stated introspection rules in nearly 60\% of cases on objects with strong color priors. Human participants remain faithful to their stated rules, with any apparent violations being explained by a well-documented tendency to overestimate color coverage. In contrast, we find that VLMs are excellent estimators of color coverage, yet blatantly contradict their own reasoning in their final responses. Across all models and strategies for eliciting introspective rules, world-knowledge priors systematically degrade faithfulness in ways that do not mirror human cognition. Our findings challenge the view that VLM reasoning failures are difficulty-driven and suggest that VLM introspective self-knowledge is miscalibrated, with direct implications for high-stakes deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2604_06422
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
Nemitz, Jonathan
Eickhoff, Carsten
Li, Junyi Jessy
Mahowald, Kyle
Golovanevsky, Michal
Rudman, William
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
I.2.7; I.4.m; I.4.8; I.2.0
Understanding when Vision-Language Models (VLMs) will behave unexpectedly, whether models can reliably predict their own behavior, and if models adhere to their introspective reasoning are central challenges for trustworthy deployment. To study this, we introduce the Graded Color Attribution (GCA) dataset, a controlled benchmark designed to elicit decision rules and evaluate participant faithfulness to these rules. GCA consists of line drawings that vary pixel-level color coverage across three conditions: world-knowledge recolorings, counterfactual recolorings, and shapes with no color priors. Using GCA, both VLMs and human participants establish a threshold: the minimum percentage of pixels of a given color an object must have to receive that color label. We then compare these rules with their subsequent color attribution decisions. Our findings reveal that models systematically violate their own introspective rules. For example, GPT-5-mini violates its stated introspection rules in nearly 60\% of cases on objects with strong color priors. Human participants remain faithful to their stated rules, with any apparent violations being explained by a well-documented tendency to overestimate color coverage. In contrast, we find that VLMs are excellent estimators of color coverage, yet blatantly contradict their own reasoning in their final responses. Across all models and strategies for eliciting introspective rules, world-knowledge priors systematically degrade faithfulness in ways that do not mirror human cognition. Our findings challenge the view that VLM reasoning failures are difficulty-driven and suggest that VLM introspective self-knowledge is miscalibrated, with direct implications for high-stakes deployment.
title When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
I.2.7; I.4.m; I.4.8; I.2.0
url https://arxiv.org/abs/2604.06422