A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sathe, Ashutosh, Jain, Prachi, Sitaram, Sunayana
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914836235943936
author Sathe, Ashutosh
Jain, Prachi
Sitaram, Sunayana
author_facet Sathe, Ashutosh
Jain, Prachi
Sitaram, Sunayana
contents Vision-language models (VLMs) have gained widespread adoption in both industry and academia. In this study, we propose a unified framework for systematically evaluating gender, race, and age biases in VLMs with respect to professions. Our evaluation encompasses all supported inference modes of the recent VLMs, including image-to-text, text-to-text, text-to-image, and image-to-image. Additionally, we propose an automated pipeline to generate high-quality synthetic datasets that intentionally conceal gender, race, and age information across different professional domains, both in generated text and images. The dataset includes action-based descriptions of each profession and serves as a benchmark for evaluating societal biases in vision-language models (VLMs). In our comparative analysis of widely used VLMs, we have identified that varying input-output modalities lead to discernible differences in bias magnitudes and directions. Additionally, we find that VLM models exhibit distinct biases across different bias attributes we investigated. We hope our work will help guide future progress in improving VLMs to learn socially unbiased representations. We will release our data and code.
format Preprint
id arxiv_https___arxiv_org_abs_2402_13636
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models
Sathe, Ashutosh
Jain, Prachi
Sitaram, Sunayana
Computer Vision and Pattern Recognition
Computation and Language
Computers and Society
Vision-language models (VLMs) have gained widespread adoption in both industry and academia. In this study, we propose a unified framework for systematically evaluating gender, race, and age biases in VLMs with respect to professions. Our evaluation encompasses all supported inference modes of the recent VLMs, including image-to-text, text-to-text, text-to-image, and image-to-image. Additionally, we propose an automated pipeline to generate high-quality synthetic datasets that intentionally conceal gender, race, and age information across different professional domains, both in generated text and images. The dataset includes action-based descriptions of each profession and serves as a benchmark for evaluating societal biases in vision-language models (VLMs). In our comparative analysis of widely used VLMs, we have identified that varying input-output modalities lead to discernible differences in bias magnitudes and directions. Additionally, we find that VLM models exhibit distinct biases across different bias attributes we investigated. We hope our work will help guide future progress in improving VLMs to learn socially unbiased representations. We will release our data and code.
title A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
Computers and Society
url https://arxiv.org/abs/2402.13636