Bias in Gender Bias Benchmarks: How Spurious Features Distort Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Hirota, Yusuke, Hachiuma, Ryo, Li, Boyi, Lu, Ximing, Boone, Michael Ross, Ivanovic, Boris, Choi, Yejin, Pavone, Marco, Wang, Yu-Chiang Frank, Garcia, Noa, Nakashima, Yuta, Yang, Chao-Han Huck |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LOTUS: A Leaderboard for Detailed Image Captioning from Quality to Societal Bias and User Preferences
by: Hirota, Yusuke, et al.
Published: (2025)
by: Hirota, Yusuke, et al.
Published: (2025)
From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment
by: Hirota, Yusuke, et al.
Published: (2024)
by: Hirota, Yusuke, et al.
Published: (2024)
From Global to Local: Social Bias Transfer in CLIP
by: Ramos, Ryan, et al.
Published: (2025)
by: Ramos, Ryan, et al.
Published: (2025)
Would Deep Generative Models Amplify Bias in Future Models?
by: Chen, Tianwei, et al.
Published: (2024)
by: Chen, Tianwei, et al.
Published: (2024)
Stable Diffusion Exposed: Gender Bias from Prompt to Image
by: Wu, Yankun, et al.
Published: (2023)
by: Wu, Yankun, et al.
Published: (2023)
Gender Bias Evaluation in Text-to-image Generation: A Survey
by: Wu, Yankun, et al.
Published: (2024)
by: Wu, Yankun, et al.
Published: (2024)
SANER: Annotation-free Societal Attribute Neutralizer for Debiasing CLIP
by: Hirota, Yusuke, et al.
Published: (2024)
by: Hirota, Yusuke, et al.
Published: (2024)
Resampled Datasets Are Not Enough: Mitigating Societal Bias Beyond Single Attributes
by: Hirota, Yusuke, et al.
Published: (2024)
by: Hirota, Yusuke, et al.
Published: (2024)
EMMA: Concept Erasure Benchmark with Comprehensive Semantic Metrics and Diverse Categories
by: Wei, Lu, et al.
Published: (2025)
by: Wei, Lu, et al.
Published: (2025)
Benchmarking Spurious Bias in Few-Shot Image Classifiers
by: Zheng, Guangtao, et al.
Published: (2024)
by: Zheng, Guangtao, et al.
Published: (2024)
Benchmarking Fairness in Spiking Neural Networks: Data Bias, Spurious Features, and Hardware Effects
by: He, Hudi, et al.
Published: (2026)
by: He, Hudi, et al.
Published: (2026)
Interpretable Debiasing of Vision-Language Models for Social Fairness
by: An, Na Min, et al.
Published: (2026)
by: An, Na Min, et al.
Published: (2026)
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
by: Sclar, Melanie, et al.
Published: (2023)
by: Sclar, Melanie, et al.
Published: (2023)
Accelerating Online Mapping and Behavior Prediction via Direct BEV Feature Attention
by: Gu, Xunjiang, et al.
Published: (2024)
by: Gu, Xunjiang, et al.
Published: (2024)
Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation
by: Ishikawa, Reina, et al.
Published: (2025)
by: Ishikawa, Reina, et al.
Published: (2025)
LoRA3D: Low-Rank Self-Calibration of 3D Geometric Foundation Models
by: Lu, Ziqi, et al.
Published: (2024)
by: Lu, Ziqi, et al.
Published: (2024)
Driving Everywhere with Large Language Model Policy Adaptation
by: Li, Boyi, et al.
Published: (2024)
by: Li, Boyi, et al.
Published: (2024)
Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
by: Lee, Byung-Kwan, et al.
Published: (2025)
by: Lee, Byung-Kwan, et al.
Published: (2025)
GenderBias-\emph{VL}: Benchmarking Gender Bias in Vision Language Models via Counterfactual Probing
by: Xiao, Yisong, et al.
Published: (2024)
by: Xiao, Yisong, et al.
Published: (2024)
Extrapolated Urban View Synthesis Benchmark
by: Han, Xiangyu, et al.
Published: (2024)
by: Han, Xiangyu, et al.
Published: (2024)
When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs
by: Wang, Shaowen, et al.
Published: (2025)
by: Wang, Shaowen, et al.
Published: (2025)
The Instinctive Bias: Spurious Images lead to Illusion in MLLMs
by: Han, Tianyang, et al.
Published: (2024)
by: Han, Tianyang, et al.
Published: (2024)
Weakly Semi-supervised Tool Detection in Minimally Invasive Surgery Videos
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
CrowdMAC: Masked Crowd Density Completion for Robust Crowd Density Forecasting
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
VIOLA: Towards Video In-Context Learning with Minimal Annotations
by: Fujii, Ryo, et al.
Published: (2026)
by: Fujii, Ryo, et al.
Published: (2026)
Towards Predicting Any Human Trajectory In Context
by: Fujii, Ryo, et al.
Published: (2025)
by: Fujii, Ryo, et al.
Published: (2025)
RealTraj: Towards Real-World Pedestrian Trajectory Forecasting
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
Tokenize the World into Object-level Knowledge to Address Long-tail Events in Autonomous Driving
by: Tian, Ran, et al.
Published: (2024)
by: Tian, Ran, et al.
Published: (2024)
Promptable Closed-loop Traffic Simulation
by: Tan, Shuhan, et al.
Published: (2024)
by: Tan, Shuhan, et al.
Published: (2024)
Benchmarking Gender and Political Bias in Large Language Models
by: Yang, Jinrui, et al.
Published: (2025)
by: Yang, Jinrui, et al.
Published: (2025)
Measuring Political Bias in Large Language Models: What Is Said and How It Is Said
by: Bang, Yejin, et al.
Published: (2024)
by: Bang, Yejin, et al.
Published: (2024)
NeuronTune: Towards Self-Guided Spurious Bias Mitigation
by: Zheng, Guangtao, et al.
Published: (2025)
by: Zheng, Guangtao, et al.
Published: (2025)
Context Matters: Auditing Gender Bias in T2I Generation through Risk-Tiered Use-Case Profiles
by: Luna, Jose, et al.
Published: (2026)
by: Luna, Jose, et al.
Published: (2026)
What is Your Favorite Gender, MLM? Gender Bias Evaluation in Multilingual Masked Language Models
by: Yu, Jeongrok, et al.
Published: (2024)
by: Yu, Jeongrok, et al.
Published: (2024)
Do Bias Benchmarks Generalise? Evidence from Voice-based Evaluation of Gender Bias in SpeechLLMs
by: Satish, Shree Harsha Bokkahalli, et al.
Published: (2025)
by: Satish, Shree Harsha Bokkahalli, et al.
Published: (2025)
Beyond Linear Bias Expansions for AbacusSummit Halos at z = 8
by: Boone, Kyle K., et al.
Published: (2026)
by: Boone, Kyle K., et al.
Published: (2026)
GG-BBQ: German Gender Bias Benchmark for Question Answering
by: Satheesh, Shalaka, et al.
Published: (2025)
by: Satheesh, Shalaka, et al.
Published: (2025)
Producing and Leveraging Online Map Uncertainty in Trajectory Prediction
by: Gu, Xunjiang, et al.
Published: (2024)
by: Gu, Xunjiang, et al.
Published: (2024)
SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal Bias
by: Ye, Wenqian, et al.
Published: (2025)
by: Ye, Wenqian, et al.
Published: (2025)
Identifying Spurious Biases Early in Training through the Lens of Simplicity Bias
by: Yang, Yu, et al.
Published: (2023)
by: Yang, Yu, et al.
Published: (2023)
Similar Items
-
LOTUS: A Leaderboard for Detailed Image Captioning from Quality to Societal Bias and User Preferences
by: Hirota, Yusuke, et al.
Published: (2025) -
From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment
by: Hirota, Yusuke, et al.
Published: (2024) -
From Global to Local: Social Bias Transfer in CLIP
by: Ramos, Ryan, et al.
Published: (2025) -
Would Deep Generative Models Amplify Bias in Future Models?
by: Chen, Tianwei, et al.
Published: (2024) -
Stable Diffusion Exposed: Gender Bias from Prompt to Image
by: Wu, Yankun, et al.
Published: (2023)