Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Bufan, Kreiss, Elisa |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GKnow: Measuring the Entanglement of Gender Bias and Factual Gender
by: Veloso, Leonor, et al.
Published: (2026)
by: Veloso, Leonor, et al.
Published: (2026)
Measuring Gender Bias in Job Title Matching for Grammatical Gender Languages
by: García-Sardiña, Laura, et al.
Published: (2025)
by: García-Sardiña, Laura, et al.
Published: (2025)
Leveraging Large Language Models to Measure Gender Representation Bias in Gendered Language Corpora
by: Derner, Erik, et al.
Published: (2024)
by: Derner, Erik, et al.
Published: (2024)
Decoding Biases: Automated Methods and LLM Judges for Gender Bias Detection in Language Models
by: Kumar, Shachi H, et al.
Published: (2024)
by: Kumar, Shachi H, et al.
Published: (2024)
When More Words Say Less: Decoupling Length and Specificity in Image Description Evaluation
by: Kapur, Rhea, et al.
Published: (2026)
by: Kapur, Rhea, et al.
Published: (2026)
VisBias: Measuring Explicit and Implicit Social Biases in Vision Language Models
by: Huang, Jen-tse, et al.
Published: (2025)
by: Huang, Jen-tse, et al.
Published: (2025)
Measurement of LLM's Philosophies of Human Nature
by: Ni, Minheng, et al.
Published: (2025)
by: Ni, Minheng, et al.
Published: (2025)
IndiBias: A Benchmark Dataset to Measure Social Biases in Language Models for Indian Context
by: Sahoo, Nihar Ranjan, et al.
Published: (2024)
by: Sahoo, Nihar Ranjan, et al.
Published: (2024)
Probability of Differentiation Reveals Brittleness of Homogeneity Bias in GPT-4
by: Lee, Messi H. J., et al.
Published: (2024)
by: Lee, Messi H. J., et al.
Published: (2024)
Analyzing Correlations Between Intrinsic and Extrinsic Bias Metrics of Static Word Embeddings With Their Measuring Biases Aligned
by: Katô, Taisei, et al.
Published: (2024)
by: Katô, Taisei, et al.
Published: (2024)
Towards Fair Rankings: Leveraging LLMs for Gender Bias Detection and Measurement
by: Mousavian, Maryam, et al.
Published: (2025)
by: Mousavian, Maryam, et al.
Published: (2025)
Measuring Opinion Bias and Sycophancy via LLM-based Persuasion
by: Nogueira, Rodrigo, et al.
Published: (2026)
by: Nogueira, Rodrigo, et al.
Published: (2026)
LLMs are Biased Teachers: Evaluating LLM Bias in Personalized Education
by: Weissburg, Iain, et al.
Published: (2024)
by: Weissburg, Iain, et al.
Published: (2024)
CommVQA: Situating Visual Question Answering in Communicative Contexts
by: Naik, Nandita Shankar, et al.
Published: (2024)
by: Naik, Nandita Shankar, et al.
Published: (2024)
From Measurement to Mitigation: Exploring the Transferability of Debiasing Approaches to Gender Bias in Maltese Language Models
by: Galea, Melanie, et al.
Published: (2025)
by: Galea, Melanie, et al.
Published: (2025)
CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
by: Verma, Arnav, et al.
Published: (2025)
by: Verma, Arnav, et al.
Published: (2025)
Measuring the Effect of Transcription Noise on Downstream Language Understanding Tasks
by: Shapira, Ori, et al.
Published: (2025)
by: Shapira, Ori, et al.
Published: (2025)
Measuring Hong Kong Massive Multi-Task Language Understanding
by: Cao, Chuxue, et al.
Published: (2025)
by: Cao, Chuxue, et al.
Published: (2025)
Trustworthy Social Bias Measurement
by: Bommasani, Rishi, et al.
Published: (2022)
by: Bommasani, Rishi, et al.
Published: (2022)
Bias Beware: The Impact of Cognitive Biases on LLM-Driven Product Recommendations
by: Filandrianos, Giorgos, et al.
Published: (2025)
by: Filandrianos, Giorgos, et al.
Published: (2025)
Measuring Stereotype and Deviation Biases in Large Language Models
by: Wang, Daniel, et al.
Published: (2025)
by: Wang, Daniel, et al.
Published: (2025)
Bias Vector: Mitigating Biases in Language Models with Task Arithmetic Approach
by: Shirafuji, Daiki, et al.
Published: (2024)
by: Shirafuji, Daiki, et al.
Published: (2024)
GenderBench: Evaluation Suite for Gender Biases in LLMs
by: Pikuliak, Matúš
Published: (2025)
by: Pikuliak, Matúš
Published: (2025)
Undesirable Biases in NLP: Addressing Challenges of Measurement
by: van der Wal, Oskar, et al.
Published: (2022)
by: van der Wal, Oskar, et al.
Published: (2022)
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance
by: Röttger, Paul, et al.
Published: (2025)
by: Röttger, Paul, et al.
Published: (2025)
Overview of the NLPCC 2025 Shared Task: Gender Bias Mitigation Challenge
by: Li, Yizhi, et al.
Published: (2025)
by: Li, Yizhi, et al.
Published: (2025)
Gender Bias in LLM-generated Interview Responses
by: Kong, Haein, et al.
Published: (2024)
by: Kong, Haein, et al.
Published: (2024)
Understanding Gender Bias in AI-Generated Product Descriptions
by: Kelly, Markelle, et al.
Published: (2025)
by: Kelly, Markelle, et al.
Published: (2025)
Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language Models
by: Kumar, Abhishek, et al.
Published: (2024)
by: Kumar, Abhishek, et al.
Published: (2024)
Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing
by: Yu, Zeping, et al.
Published: (2025)
by: Yu, Zeping, et al.
Published: (2025)
New Job, New Gender? Measuring the Social Bias in Image Generation Models
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
Bias in News Summarization: Measures, Pitfalls and Corpora
by: Steen, Julius, et al.
Published: (2023)
by: Steen, Julius, et al.
Published: (2023)
Measuring Social Biases in Masked Language Models by Proxy of Prediction Quality
by: Zalkikar, Rahul, et al.
Published: (2024)
by: Zalkikar, Rahul, et al.
Published: (2024)
Robust Evaluation Measures for Evaluating Social Biases in Masked Language Models
by: Liu, Yang
Published: (2024)
by: Liu, Yang
Published: (2024)
Colombian Waitresses y Jueces canadienses: Gender and Country Biases in Occupation Recommendations from LLMs
by: Rodríguez, Elisa Forcada, et al.
Published: (2025)
by: Rodríguez, Elisa Forcada, et al.
Published: (2025)
LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance
by: Haller, Patrick, et al.
Published: (2025)
by: Haller, Patrick, et al.
Published: (2025)
Weird Generalization is Weirdly Brittle
by: Wanner, Miriam, et al.
Published: (2026)
by: Wanner, Miriam, et al.
Published: (2026)
Brittle Minds, Fixable Activations: Understanding Belief Representations in Language Models
by: Bortoletto, Matteo, et al.
Published: (2024)
by: Bortoletto, Matteo, et al.
Published: (2024)
For the Misgendered Chinese in Gender Bias Research: Multi-Task Learning with Knowledge Distillation for Pinyin Name-Gender Prediction
by: Du, Xiaocong, et al.
Published: (2024)
by: Du, Xiaocong, et al.
Published: (2024)
Are Bias Evaluation Methods Biased ?
by: Berrayana, Lina, et al.
Published: (2025)
by: Berrayana, Lina, et al.
Published: (2025)
Similar Items
-
GKnow: Measuring the Entanglement of Gender Bias and Factual Gender
by: Veloso, Leonor, et al.
Published: (2026) -
Measuring Gender Bias in Job Title Matching for Grammatical Gender Languages
by: García-Sardiña, Laura, et al.
Published: (2025) -
Leveraging Large Language Models to Measure Gender Representation Bias in Gendered Language Corpora
by: Derner, Erik, et al.
Published: (2024) -
Decoding Biases: Automated Methods and LLM Judges for Gender Bias Detection in Language Models
by: Kumar, Shachi H, et al.
Published: (2024) -
When More Words Say Less: Decoupling Length and Specificity in Image Description Evaluation
by: Kapur, Rhea, et al.
Published: (2026)