Pinpoint Counterfactuals: Reducing social bias in foundation models via localized counterfactual generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929627500380160 |
|---|---|
| author | Sirotkin, Kirill Escudero-Viñolo, Marcos Carballeira, Pablo Maniparambil, Mayug Barata, Catarina O'Connor, Noel E. |
| author_facet | Sirotkin, Kirill Escudero-Viñolo, Marcos Carballeira, Pablo Maniparambil, Mayug Barata, Catarina O'Connor, Noel E. |
| contents | Foundation models trained on web-scraped datasets propagate societal biases to downstream tasks. While counterfactual generation enables bias analysis, existing methods introduce artifacts by modifying contextual elements like clothing and background. We present a localized counterfactual generation method that preserves image context by constraining counterfactual modifications to specific attribute-relevant regions through automated masking and guided inpainting. When applied to the Conceptual Captions dataset for creating gender counterfactuals, our method results in higher visual and semantic fidelity than state-of-the-art alternatives, while maintaining the performance of models trained using only real data on non-human-centric tasks. Models fine-tuned with our counterfactuals demonstrate measurable bias reduction across multiple metrics, including a decrease in gender classification disparity and balanced person preference scores, while preserving ImageNet zero-shot performance. The results establish a framework for creating balanced datasets that enable both accurate bias profiling and effective mitigation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_09160 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Pinpoint Counterfactuals: Reducing social bias in foundation models via localized counterfactual generation Sirotkin, Kirill Escudero-Viñolo, Marcos Carballeira, Pablo Maniparambil, Mayug Barata, Catarina O'Connor, Noel E. Computer Vision and Pattern Recognition Foundation models trained on web-scraped datasets propagate societal biases to downstream tasks. While counterfactual generation enables bias analysis, existing methods introduce artifacts by modifying contextual elements like clothing and background. We present a localized counterfactual generation method that preserves image context by constraining counterfactual modifications to specific attribute-relevant regions through automated masking and guided inpainting. When applied to the Conceptual Captions dataset for creating gender counterfactuals, our method results in higher visual and semantic fidelity than state-of-the-art alternatives, while maintaining the performance of models trained using only real data on non-human-centric tasks. Models fine-tuned with our counterfactuals demonstrate measurable bias reduction across multiple metrics, including a decrease in gender classification disparity and balanced person preference scores, while preserving ImageNet zero-shot performance. The results establish a framework for creating balanced datasets that enable both accurate bias profiling and effective mitigation. |
| title | Pinpoint Counterfactuals: Reducing social bias in foundation models via localized counterfactual generation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2412.09160 |