Debiasing Large Vision-Language Models by Ablating Protected Attribute Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ratzlaff, Neale, Olson, Matthew Lyle, Hinck, Musashi, Tseng, Shao-Yen, Lal, Vasudev, Howard, Phillip
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908096377389056
author Ratzlaff, Neale
Olson, Matthew Lyle
Hinck, Musashi
Tseng, Shao-Yen
Lal, Vasudev
Howard, Phillip
author_facet Ratzlaff, Neale
Olson, Matthew Lyle
Hinck, Musashi
Tseng, Shao-Yen
Lal, Vasudev
Howard, Phillip
contents Large Vision Language Models (LVLMs) such as LLaVA have demonstrated impressive capabilities as general-purpose chatbots that can engage in conversations about a provided input image. However, their responses are influenced by societal biases present in their training datasets, leading to undesirable differences in how the model responds when presented with images depicting people of different demographics. In this work, we propose a novel debiasing framework for LVLMs by directly ablating biased attributes during text generation to avoid generating text related to protected attributes, or even representing them internally. Our method requires no training and a relatively small amount of representative biased outputs (~1000 samples). Our experiments show that not only can we can minimize the propensity of LVLMs to generate text related to protected attributes, but we can even use synthetic data to inform the ablation while retaining captioning performance on real data such as COCO. Furthermore, we find the resulting generations from a debiased LVLM exhibit similar accuracy as a baseline biased model, showing that debiasing effects can be achieved without sacrificing model performance.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13976
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Debiasing Large Vision-Language Models by Ablating Protected Attribute Representations
Ratzlaff, Neale
Olson, Matthew Lyle
Hinck, Musashi
Tseng, Shao-Yen
Lal, Vasudev
Howard, Phillip
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Large Vision Language Models (LVLMs) such as LLaVA have demonstrated impressive capabilities as general-purpose chatbots that can engage in conversations about a provided input image. However, their responses are influenced by societal biases present in their training datasets, leading to undesirable differences in how the model responds when presented with images depicting people of different demographics. In this work, we propose a novel debiasing framework for LVLMs by directly ablating biased attributes during text generation to avoid generating text related to protected attributes, or even representing them internally. Our method requires no training and a relatively small amount of representative biased outputs (~1000 samples). Our experiments show that not only can we can minimize the propensity of LVLMs to generate text related to protected attributes, but we can even use synthetic data to inform the ablation while retaining captioning performance on real data such as COCO. Furthermore, we find the resulting generations from a debiased LVLM exhibit similar accuracy as a baseline biased model, showing that debiasing effects can be achieved without sacrificing model performance.
title Debiasing Large Vision-Language Models by Ablating Protected Attribute Representations
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
url https://arxiv.org/abs/2410.13976