Vision Language Models are Biased

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vo, An, Nguyen, Khai-Nguyen, Taesiri, Mohammad Reza, Dang, Vy Tuong, Nguyen, Anh Totti, Kim, Daeyoung
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913043512819712
author Vo, An
Nguyen, Khai-Nguyen
Taesiri, Mohammad Reza
Dang, Vy Tuong
Nguyen, Anh Totti
Kim, Daeyoung
author_facet Vo, An
Nguyen, Khai-Nguyen
Taesiri, Mohammad Reza
Dang, Vy Tuong
Nguyen, Anh Totti
Kim, Daeyoung
contents Large language models (LLMs) memorize a vast amount of prior knowledge from the Internet that helps them on downstream tasks but also may notoriously sway their outputs towards wrong or biased answers. In this work, we test how the knowledge about popular subjects hurt the accuracy of vision language models (VLMs) on standard, objective visual tasks of counting and identification. We find that state-of-the-art VLMs are strongly biased (e.g., unable to recognize the 4th stripe has been added to a 3-stripe Adidas logo) scoring an average of 17.05% accuracy in counting (e.g., counting stripes in an Adidas-like logo) across 7 diverse domains from animals, logos, chess, board games, optical illusions, to patterned grids. Removing image backgrounds nearly doubles accuracy (21.09 percentage points), revealing that contextual visual cues trigger these biased responses. Further analysis of VLMs' reasoning patterns shows that counting accuracy initially rises with thinking tokens, reaching ~40%, before declining with excessive reasoning. Our work presents an interesting failure mode in VLMs and a human-supervised automated framework for testing VLM biases. Code and data are available at: vlmsarebiased.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23941
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Vision Language Models are Biased
Vo, An
Nguyen, Khai-Nguyen
Taesiri, Mohammad Reza
Dang, Vy Tuong
Nguyen, Anh Totti
Kim, Daeyoung
Machine Learning
Computer Vision and Pattern Recognition
Large language models (LLMs) memorize a vast amount of prior knowledge from the Internet that helps them on downstream tasks but also may notoriously sway their outputs towards wrong or biased answers. In this work, we test how the knowledge about popular subjects hurt the accuracy of vision language models (VLMs) on standard, objective visual tasks of counting and identification. We find that state-of-the-art VLMs are strongly biased (e.g., unable to recognize the 4th stripe has been added to a 3-stripe Adidas logo) scoring an average of 17.05% accuracy in counting (e.g., counting stripes in an Adidas-like logo) across 7 diverse domains from animals, logos, chess, board games, optical illusions, to patterned grids. Removing image backgrounds nearly doubles accuracy (21.09 percentage points), revealing that contextual visual cues trigger these biased responses. Further analysis of VLMs' reasoning patterns shows that counting accuracy initially rises with thinking tokens, reaching ~40%, before declining with excessive reasoning. Our work presents an interesting failure mode in VLMs and a human-supervised automated framework for testing VLM biases. Code and data are available at: vlmsarebiased.github.io.
title Vision Language Models are Biased
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.23941