GG-BBQ: German Gender Bias Benchmark for Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Satheesh, Shalaka, Klug, Katrin, Beckh, Katharina, Allende-Cid, Héctor, Houben, Sebastian, Hassan, Teena |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Continual Pre-training Bridge the Performance Gap between General-purpose and Specialized Language Models in the Medical Domain?
by: Doll, Niclas, et al.
Published: (2026)
by: Doll, Niclas, et al.
Published: (2026)
PakBBQ: A Culturally Adapted Bias Benchmark for QA
by: Hashmat, Abdullah, et al.
Published: (2025)
by: Hashmat, Abdullah, et al.
Published: (2025)
EsBBQ and CaBBQ: The Spanish and Catalan Bias Benchmarks for Question Answering
by: Ruiz-Fernández, Valle, et al.
Published: (2025)
by: Ruiz-Fernández, Valle, et al.
Published: (2025)
KoBBQ: Korean Bias Benchmark for Question Answering
by: Jin, Jiho, et al.
Published: (2023)
by: Jin, Jiho, et al.
Published: (2023)
BharatBBQ: A Multilingual Bias Benchmark for Question Answering in the Indian Context
by: Tomar, Aditya, et al.
Published: (2025)
by: Tomar, Aditya, et al.
Published: (2025)
Mitigating Bias for Question Answering Models by Tracking Bias Influence
by: Ma, Mingyu Derek, et al.
Published: (2023)
by: Ma, Mingyu Derek, et al.
Published: (2023)
Social Bias in Popular Question-Answering Benchmarks
by: Kraft, Angelie, et al.
Published: (2025)
by: Kraft, Angelie, et al.
Published: (2025)
Robust Bias Evaluation with FilBBQ: A Filipino Bias Benchmark for Question-Answering Language Models
by: Gamboa, Lance Calvin Lim, et al.
Published: (2026)
by: Gamboa, Lance Calvin Lim, et al.
Published: (2026)
How Prevalent is Gender Bias in ChatGPT? -- Exploring German and English ChatGPT Responses
by: Urchs, Stefanie, et al.
Published: (2023)
by: Urchs, Stefanie, et al.
Published: (2023)
Exploring the Linear Subspace Hypothesis in Gender Bias Mitigation
by: Vargas, Francisco, et al.
Published: (2020)
by: Vargas, Francisco, et al.
Published: (2020)
Towards Unsupervised Question Answering System with Multi-level Summarization for Legal Text
by: Prabhu, M Manvith, et al.
Published: (2024)
by: Prabhu, M Manvith, et al.
Published: (2024)
Complete Evidence Extraction with Model Ensembles: A Case Study on Medical Coding
by: Beckh, Katharina, et al.
Published: (2025)
by: Beckh, Katharina, et al.
Published: (2025)
SyllabusQA: A Course Logistics Question Answering Dataset
by: Fernandez, Nigel, et al.
Published: (2024)
by: Fernandez, Nigel, et al.
Published: (2024)
Beyond Prompting: An Efficient Embedding Framework for Open-Domain Question Answering
by: Hu, Zhanghao, et al.
Published: (2025)
by: Hu, Zhanghao, et al.
Published: (2025)
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
by: Majumdar, Ayan, et al.
Published: (2025)
by: Majumdar, Ayan, et al.
Published: (2025)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
Multilingual Text-to-Image Generation Magnifies Gender Stereotypes and Prompt Engineering May Not Help You
by: Friedrich, Felix, et al.
Published: (2024)
by: Friedrich, Felix, et al.
Published: (2024)
What's in a Name? Auditing Large Language Models for Race and Gender Bias
by: Salinas, Alejandro, et al.
Published: (2024)
by: Salinas, Alejandro, et al.
Published: (2024)
Asking For It: Question-Answering for Predicting Rule Infractions in Online Content Moderation
by: Samory, Mattia, et al.
Published: (2025)
by: Samory, Mattia, et al.
Published: (2025)
Leveraging Large Language Models to Measure Gender Representation Bias in Gendered Language Corpora
by: Derner, Erik, et al.
Published: (2024)
by: Derner, Erik, et al.
Published: (2024)
The Anatomy of Evidence: An Investigation Into Explainable ICD Coding
by: Beckh, Katharina, et al.
Published: (2025)
by: Beckh, Katharina, et al.
Published: (2025)
Gender Bias in Emotion Recognition by Large Language Models
by: Herbert, Maureen, et al.
Published: (2025)
by: Herbert, Maureen, et al.
Published: (2025)
JobFair: A Framework for Benchmarking Gender Hiring Bias in Large Language Models
by: Wang, Ze, et al.
Published: (2024)
by: Wang, Ze, et al.
Published: (2024)
PediatricsMQA: a Multi-modal Pediatrics Question Answering Benchmark
by: Bahaj, Adil, et al.
Published: (2025)
by: Bahaj, Adil, et al.
Published: (2025)
GECOBench: A Gender-Controlled Text Dataset and Benchmark for Quantifying Biases in Explanations
by: Wilming, Rick, et al.
Published: (2024)
by: Wilming, Rick, et al.
Published: (2024)
CAIRNS: Balancing Readability and Scientific Accuracy in Climate Adaptation Question Answering
by: Kong, Liangji, et al.
Published: (2025)
by: Kong, Liangji, et al.
Published: (2025)
Questionable practices in machine learning
by: Leech, Gavin, et al.
Published: (2024)
by: Leech, Gavin, et al.
Published: (2024)
Reflecting in the Reflection: Integrating a Socratic Questioning Framework into Automated AI-Based Question Generation
by: Holub, Ondřej, et al.
Published: (2026)
by: Holub, Ondřej, et al.
Published: (2026)
Gender Representation and Bias in Indian Civil Service Mock Interviews
by: Banerjee, Somonnoy, et al.
Published: (2024)
by: Banerjee, Somonnoy, et al.
Published: (2024)
Question-Answering (QA) Model for a Personalized Learning Assistant for Arabic Language
by: Sammoudi, Mohammad, et al.
Published: (2024)
by: Sammoudi, Mohammad, et al.
Published: (2024)
SUKHSANDESH: An Avatar Therapeutic Question Answering Platform for Sexual Education in Rural India
by: Singh, Salam Michael, et al.
Published: (2024)
by: Singh, Salam Michael, et al.
Published: (2024)
MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing
by: Agarwal, Siddhant, et al.
Published: (2024)
by: Agarwal, Siddhant, et al.
Published: (2024)
PopBERT. Detecting populism and its host ideologies in the German Bundestag
by: Erhard, L., et al.
Published: (2023)
by: Erhard, L., et al.
Published: (2023)
DSO: Direct Steering Optimization for Bias Mitigation
by: Paes, Lucas Monteiro, et al.
Published: (2025)
by: Paes, Lucas Monteiro, et al.
Published: (2025)
Llms, Virtual Users, and Bias: Predicting Any Survey Question Without Human Data
by: Sinacola, Enzo, et al.
Published: (2025)
by: Sinacola, Enzo, et al.
Published: (2025)
For the Misgendered Chinese in Gender Bias Research: Multi-Task Learning with Knowledge Distillation for Pinyin Name-Gender Prediction
by: Du, Xiaocong, et al.
Published: (2024)
by: Du, Xiaocong, et al.
Published: (2024)
Gender Bias Detection in Court Decisions: A Brazilian Case Study
by: Benatti, Raysa, et al.
Published: (2024)
by: Benatti, Raysa, et al.
Published: (2024)
Assessing Modality Bias in Video Question Answering Benchmarks with Multimodal Large Language Models
by: Park, Jean, et al.
Published: (2024)
by: Park, Jean, et al.
Published: (2024)
Addressing Both Statistical and Causal Gender Fairness in NLP Models
by: Chen, Hannah, et al.
Published: (2024)
by: Chen, Hannah, et al.
Published: (2024)
Fair Representation in Parliamentary Summaries: Measuring and Mitigating Inclusion Bias
by: Cunningham, Eoghan, et al.
Published: (2025)
by: Cunningham, Eoghan, et al.
Published: (2025)
Similar Items
-
Can Continual Pre-training Bridge the Performance Gap between General-purpose and Specialized Language Models in the Medical Domain?
by: Doll, Niclas, et al.
Published: (2026) -
PakBBQ: A Culturally Adapted Bias Benchmark for QA
by: Hashmat, Abdullah, et al.
Published: (2025) -
EsBBQ and CaBBQ: The Spanish and Catalan Bias Benchmarks for Question Answering
by: Ruiz-Fernández, Valle, et al.
Published: (2025) -
KoBBQ: Korean Bias Benchmark for Question Answering
by: Jin, Jiho, et al.
Published: (2023) -
BharatBBQ: A Multilingual Bias Benchmark for Question Answering in the Indian Context
by: Tomar, Aditya, et al.
Published: (2025)