Social Bias in Popular Question-Answering Benchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | Kraft, Angelie, Simon, Judith, Schimmler, Sonja |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Lifecycle of "Facts": A Survey of Social Bias in Knowledge Graphs
by: Kraft, Angelie, et al.
Published: (2022)
by: Kraft, Angelie, et al.
Published: (2022)
Mitigating Bias for Question Answering Models by Tracking Bias Influence
by: Ma, Mingyu Derek, et al.
Published: (2023)
by: Ma, Mingyu Derek, et al.
Published: (2023)
PediatricsMQA: a Multi-modal Pediatrics Question Answering Benchmark
by: Bahaj, Adil, et al.
Published: (2025)
by: Bahaj, Adil, et al.
Published: (2025)
KoBBQ: Korean Bias Benchmark for Question Answering
by: Jin, Jiho, et al.
Published: (2023)
by: Jin, Jiho, et al.
Published: (2023)
From Individuals to Interactions: Benchmarking Gender Bias in Multimodal Large Language Models from the Lens of Social Relationship
by: Xu, Yue, et al.
Published: (2025)
by: Xu, Yue, et al.
Published: (2025)
Reporting and Analysing the Environmental Impact of Language Models on the Example of Commonsense Question Answering with External Knowledge
by: Usmanova, Aida, et al.
Published: (2024)
by: Usmanova, Aida, et al.
Published: (2024)
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
by: Haider, Batool, et al.
Published: (2025)
by: Haider, Batool, et al.
Published: (2025)
AccessEval: Benchmarking Disability Bias in Large Language Models
by: Panda, Srikant, et al.
Published: (2025)
by: Panda, Srikant, et al.
Published: (2025)
DeCAP: Context-Adaptive Prompt Generation for Debiasing Zero-shot Question Answering in Large Language Models
by: Bae, Suyoung, et al.
Published: (2025)
by: Bae, Suyoung, et al.
Published: (2025)
Beyond Prompting: An Efficient Embedding Framework for Open-Domain Question Answering
by: Hu, Zhanghao, et al.
Published: (2025)
by: Hu, Zhanghao, et al.
Published: (2025)
JobFair: A Framework for Benchmarking Gender Hiring Bias in Large Language Models
by: Wang, Ze, et al.
Published: (2024)
by: Wang, Ze, et al.
Published: (2024)
Rethinking LLM Bias Probing Using Lessons from the Social Sciences
by: Morehouse, Kirsten N., et al.
Published: (2025)
by: Morehouse, Kirsten N., et al.
Published: (2025)
Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language Models
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
Ask LLMs Directly, "What shapes your bias?": Measuring Social Bias in Large Language Models
by: Shin, Jisu, et al.
Published: (2024)
by: Shin, Jisu, et al.
Published: (2024)
GG-BBQ: German Gender Bias Benchmark for Question Answering
by: Satheesh, Shalaka, et al.
Published: (2025)
by: Satheesh, Shalaka, et al.
Published: (2025)
Dual-Metric Evaluation of Social Bias in Large Language Models: Evidence from an Underrepresented Nepali Cultural Context
by: Pandey, Ashish, et al.
Published: (2026)
by: Pandey, Ashish, et al.
Published: (2026)
PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media
by: Kachwala, Zoher, et al.
Published: (2026)
by: Kachwala, Zoher, et al.
Published: (2026)
PakBBQ: A Culturally Adapted Bias Benchmark for QA
by: Hashmat, Abdullah, et al.
Published: (2025)
by: Hashmat, Abdullah, et al.
Published: (2025)
Automated Item Neutralization for Non-Cognitive Scales: A Large Language Model Approach to Reducing Social-Desirability Bias
by: Wu, Sirui, et al.
Published: (2025)
by: Wu, Sirui, et al.
Published: (2025)
Red Lines and Grey Zones in the Fog of War: Benchmarking Legal Risk, Moral Harm, and Regional Bias in Large Language Model Military Decision-Making
by: Drinkall, Toby
Published: (2025)
by: Drinkall, Toby
Published: (2025)
Different Bias Under Different Criteria: Assessing Bias in LLMs with a Fact-Based Approach
by: Ko, Changgeon, et al.
Published: (2024)
by: Ko, Changgeon, et al.
Published: (2024)
Understanding Social Support Needs in Questions: A Hybrid Approach Integrating Semi-Supervised Learning and LLM-based Data Augmentation
by: Kuang, Junwei, et al.
Published: (2025)
by: Kuang, Junwei, et al.
Published: (2025)
Asking For It: Question-Answering for Predicting Rule Infractions in Online Content Moderation
by: Samory, Mattia, et al.
Published: (2025)
by: Samory, Mattia, et al.
Published: (2025)
No Free Lunch in Language Model Bias Mitigation? Targeted Bias Reduction Can Exacerbate Unmitigated LLM Biases
by: Chand, Shireen, et al.
Published: (2025)
by: Chand, Shireen, et al.
Published: (2025)
White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMs
by: Wan, Yixin, et al.
Published: (2024)
by: Wan, Yixin, et al.
Published: (2024)
On the Credibility of Evaluating LLMs using Survey Questions
by: Libovický, Jindřich
Published: (2026)
by: Libovický, Jindřich
Published: (2026)
Cross-Language Bias Examination in Large Language Models
by: Liang, Yuxuan, et al.
Published: (2025)
by: Liang, Yuxuan, et al.
Published: (2025)
EtiCor++: Towards Understanding Etiquettical Bias in LLMs
by: Dwivedi, Ashutosh, et al.
Published: (2025)
by: Dwivedi, Ashutosh, et al.
Published: (2025)
Perceived Political Bias in LLMs Reduces Persuasive Abilities
by: DiGiuseppe, Matthew, et al.
Published: (2026)
by: DiGiuseppe, Matthew, et al.
Published: (2026)
Assessing Modality Bias in Video Question Answering Benchmarks with Multimodal Large Language Models
by: Park, Jean, et al.
Published: (2024)
by: Park, Jean, et al.
Published: (2024)
Mitigating Gender Bias via Fostering Exploratory Thinking in LLMs
by: Wei, Kangda, et al.
Published: (2025)
by: Wei, Kangda, et al.
Published: (2025)
Artificial Intelligence Bias on English Language Learners in Automatic Scoring
by: Guo, Shuchen, et al.
Published: (2025)
by: Guo, Shuchen, et al.
Published: (2025)
The Ethical Risks of Analyzing Crisis Events on Social Media with Machine Learning
by: Kraft, Angelie, et al.
Published: (2022)
by: Kraft, Angelie, et al.
Published: (2022)
Widespread Gender and Pronoun Bias in Moral Judgments Across LLMs
by: Fernandes, Gustavo Lúcius, et al.
Published: (2026)
by: Fernandes, Gustavo Lúcius, et al.
Published: (2026)
Gender Bias in Machine Translation and The Era of Large Language Models
by: Vanmassenhove, Eva
Published: (2024)
by: Vanmassenhove, Eva
Published: (2024)
A Benchmark for Long-Form Medical Question Answering
by: Hosseini, Pedram, et al.
Published: (2024)
by: Hosseini, Pedram, et al.
Published: (2024)
Sacred or Synthetic? Evaluating LLM Reliability and Abstention for Religious Questions
by: Atif, Farah, et al.
Published: (2025)
by: Atif, Farah, et al.
Published: (2025)
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race
by: Sun, Lihao, et al.
Published: (2025)
by: Sun, Lihao, et al.
Published: (2025)
InsideOut: Measuring and Mitigating Insider-Outsider Bias in Interview Script Generation
by: Wan, Yixin, et al.
Published: (2025)
by: Wan, Yixin, et al.
Published: (2025)
Similar Items
-
The Lifecycle of "Facts": A Survey of Social Bias in Knowledge Graphs
by: Kraft, Angelie, et al.
Published: (2022) -
Mitigating Bias for Question Answering Models by Tracking Bias Influence
by: Ma, Mingyu Derek, et al.
Published: (2023) -
PediatricsMQA: a Multi-modal Pediatrics Question Answering Benchmark
by: Bahaj, Adil, et al.
Published: (2025) -
KoBBQ: Korean Bias Benchmark for Question Answering
by: Jin, Jiho, et al.
Published: (2023) -
From Individuals to Interactions: Benchmarking Gender Bias in Multimodal Large Language Models from the Lens of Social Relationship
by: Xu, Yue, et al.
Published: (2025)