Pinpoint Counterfactuals: Reducing social bias in foundation models via localized counterfactual generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sirotkin, Kirill, Escudero-Viñolo, Marcos, Carballeira, Pablo, Maniparambil, Mayug, Barata, Catarina, O'Connor, Noel E.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929627500380160
author Sirotkin, Kirill
Escudero-Viñolo, Marcos
Carballeira, Pablo
Maniparambil, Mayug
Barata, Catarina
O'Connor, Noel E.
author_facet Sirotkin, Kirill
Escudero-Viñolo, Marcos
Carballeira, Pablo
Maniparambil, Mayug
Barata, Catarina
O'Connor, Noel E.
contents Foundation models trained on web-scraped datasets propagate societal biases to downstream tasks. While counterfactual generation enables bias analysis, existing methods introduce artifacts by modifying contextual elements like clothing and background. We present a localized counterfactual generation method that preserves image context by constraining counterfactual modifications to specific attribute-relevant regions through automated masking and guided inpainting. When applied to the Conceptual Captions dataset for creating gender counterfactuals, our method results in higher visual and semantic fidelity than state-of-the-art alternatives, while maintaining the performance of models trained using only real data on non-human-centric tasks. Models fine-tuned with our counterfactuals demonstrate measurable bias reduction across multiple metrics, including a decrease in gender classification disparity and balanced person preference scores, while preserving ImageNet zero-shot performance. The results establish a framework for creating balanced datasets that enable both accurate bias profiling and effective mitigation.
format Preprint
id arxiv_https___arxiv_org_abs_2412_09160
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Pinpoint Counterfactuals: Reducing social bias in foundation models via localized counterfactual generation
Sirotkin, Kirill
Escudero-Viñolo, Marcos
Carballeira, Pablo
Maniparambil, Mayug
Barata, Catarina
O'Connor, Noel E.
Computer Vision and Pattern Recognition
Foundation models trained on web-scraped datasets propagate societal biases to downstream tasks. While counterfactual generation enables bias analysis, existing methods introduce artifacts by modifying contextual elements like clothing and background. We present a localized counterfactual generation method that preserves image context by constraining counterfactual modifications to specific attribute-relevant regions through automated masking and guided inpainting. When applied to the Conceptual Captions dataset for creating gender counterfactuals, our method results in higher visual and semantic fidelity than state-of-the-art alternatives, while maintaining the performance of models trained using only real data on non-human-centric tasks. Models fine-tuned with our counterfactuals demonstrate measurable bias reduction across multiple metrics, including a decrease in gender classification disparity and balanced person preference scores, while preserving ImageNet zero-shot performance. The results establish a framework for creating balanced datasets that enable both accurate bias profiling and effective mitigation.
title Pinpoint Counterfactuals: Reducing social bias in foundation models via localized counterfactual generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.09160