Bias in Gender Bias Benchmarks: How Spurious Features Distort Evaluation
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866914077207429120 |
|---|---|
| author | Hirota, Yusuke Hachiuma, Ryo Li, Boyi Lu, Ximing Boone, Michael Ross Ivanovic, Boris Choi, Yejin Pavone, Marco Wang, Yu-Chiang Frank Garcia, Noa Nakashima, Yuta Yang, Chao-Han Huck |
| author_facet | Hirota, Yusuke Hachiuma, Ryo Li, Boyi Lu, Ximing Boone, Michael Ross Ivanovic, Boris Choi, Yejin Pavone, Marco Wang, Yu-Chiang Frank Garcia, Noa Nakashima, Yuta Yang, Chao-Han Huck |
| contents | Gender bias in vision-language foundation models (VLMs) raises concerns about their safe deployment and is typically evaluated using benchmarks with gender annotations on real-world images. However, as these benchmarks often contain spurious correlations between gender and non-gender features, such as objects and backgrounds, we identify a critical oversight in gender bias evaluation: Do spurious features distort gender bias evaluation? To address this question, we systematically perturb non-gender features across four widely used benchmarks (COCO-gender, FACET, MIAP, and PHASE) and various VLMs to quantify their impact on bias evaluation. Our findings reveal that even minimal perturbations, such as masking just 10% of objects or weakly blurring backgrounds, can dramatically alter bias scores, shifting metrics by up to 175% in generative VLMs and 43% in CLIP variants. This suggests that current bias evaluations often reflect model responses to spurious features rather than gender bias, undermining their reliability. Since creating spurious feature-free benchmarks is fundamentally challenging, we recommend reporting bias metrics alongside feature-sensitivity measurements to enable a more reliable bias assessment. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_07596 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Bias in Gender Bias Benchmarks: How Spurious Features Distort Evaluation Hirota, Yusuke Hachiuma, Ryo Li, Boyi Lu, Ximing Boone, Michael Ross Ivanovic, Boris Choi, Yejin Pavone, Marco Wang, Yu-Chiang Frank Garcia, Noa Nakashima, Yuta Yang, Chao-Han Huck Computer Vision and Pattern Recognition Gender bias in vision-language foundation models (VLMs) raises concerns about their safe deployment and is typically evaluated using benchmarks with gender annotations on real-world images. However, as these benchmarks often contain spurious correlations between gender and non-gender features, such as objects and backgrounds, we identify a critical oversight in gender bias evaluation: Do spurious features distort gender bias evaluation? To address this question, we systematically perturb non-gender features across four widely used benchmarks (COCO-gender, FACET, MIAP, and PHASE) and various VLMs to quantify their impact on bias evaluation. Our findings reveal that even minimal perturbations, such as masking just 10% of objects or weakly blurring backgrounds, can dramatically alter bias scores, shifting metrics by up to 175% in generative VLMs and 43% in CLIP variants. This suggests that current bias evaluations often reflect model responses to spurious features rather than gender bias, undermining their reliability. Since creating spurious feature-free benchmarks is fundamentally challenging, we recommend reporting bias metrics alongside feature-sensitivity measurements to enable a more reliable bias assessment. |
| title | Bias in Gender Bias Benchmarks: How Spurious Features Distort Evaluation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2509.07596 |