Robust Bias Evaluation with FilBBQ: A Filipino Bias Benchmark for Question-Answering Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gamboa, Lance Calvin Lim, Feng, Yue, Lee, Mark |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Bias Attribution in Filipino Language Models: Extending a Bias Interpretability Metric for Application on Agglutinative Languages
von: Gamboa, Lance Calvin Lim, et al.
Veröffentlicht: (2025)
von: Gamboa, Lance Calvin Lim, et al.
Veröffentlicht: (2025)
Filipino Benchmarks for Measuring Sexist and Homophobic Bias in Multilingual Language Models from Southeast Asia
von: Gamboa, Lance Calvin Lim, et al.
Veröffentlicht: (2024)
von: Gamboa, Lance Calvin Lim, et al.
Veröffentlicht: (2024)
Social Bias in Multilingual Language Models: A Survey
von: Gamboa, Lance Calvin Lim, et al.
Veröffentlicht: (2025)
von: Gamboa, Lance Calvin Lim, et al.
Veröffentlicht: (2025)
KoBBQ: Korean Bias Benchmark for Question Answering
von: Jin, Jiho, et al.
Veröffentlicht: (2023)
von: Jin, Jiho, et al.
Veröffentlicht: (2023)
A Novel Interpretability Metric for Explaining Bias in Language Models: Applications on Multilingual Models from Southeast Asia
von: Gamboa, Lance Calvin Lim, et al.
Veröffentlicht: (2024)
von: Gamboa, Lance Calvin Lim, et al.
Veröffentlicht: (2024)
EsBBQ and CaBBQ: The Spanish and Catalan Bias Benchmarks for Question Answering
von: Ruiz-Fernández, Valle, et al.
Veröffentlicht: (2025)
von: Ruiz-Fernández, Valle, et al.
Veröffentlicht: (2025)
GG-BBQ: German Gender Bias Benchmark for Question Answering
von: Satheesh, Shalaka, et al.
Veröffentlicht: (2025)
von: Satheesh, Shalaka, et al.
Veröffentlicht: (2025)
BharatBBQ: A Multilingual Bias Benchmark for Question Answering in the Indian Context
von: Tomar, Aditya, et al.
Veröffentlicht: (2025)
von: Tomar, Aditya, et al.
Veröffentlicht: (2025)
PakBBQ: A Culturally Adapted Bias Benchmark for QA
von: Hashmat, Abdullah, et al.
Veröffentlicht: (2025)
von: Hashmat, Abdullah, et al.
Veröffentlicht: (2025)
FilBench: Can LLMs Understand and Generate Filipino?
von: Miranda, Lester James V., et al.
Veröffentlicht: (2025)
von: Miranda, Lester James V., et al.
Veröffentlicht: (2025)
ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues
von: Vedula, Bhaskara Hanuma, et al.
Veröffentlicht: (2026)
von: Vedula, Bhaskara Hanuma, et al.
Veröffentlicht: (2026)
VoiceBBQ: Investigating Effect of Content and Acoustics in Social Bias of Spoken Language Model
von: Choi, Junhyuk, et al.
Veröffentlicht: (2025)
von: Choi, Junhyuk, et al.
Veröffentlicht: (2025)
Social Bias in Popular Question-Answering Benchmarks
von: Kraft, Angelie, et al.
Veröffentlicht: (2025)
von: Kraft, Angelie, et al.
Veröffentlicht: (2025)
BioPulse-QA: A Dynamic Biomedical Question-Answering Benchmark for Evaluating Factuality, Robustness, and Bias in Large Language Models
von: Bhattarai, Kriti, et al.
Veröffentlicht: (2026)
von: Bhattarai, Kriti, et al.
Veröffentlicht: (2026)
Assessing Modality Bias in Video Question Answering Benchmarks with Multimodal Large Language Models
von: Park, Jean, et al.
Veröffentlicht: (2024)
von: Park, Jean, et al.
Veröffentlicht: (2024)
Mitigating Bias for Question Answering Models by Tracking Bias Influence
von: Ma, Mingyu Derek, et al.
Veröffentlicht: (2023)
von: Ma, Mingyu Derek, et al.
Veröffentlicht: (2023)
More Bias, Less Bias: BiasPrompting for Enhanced Multiple-Choice Question Answering
von: Vu, Duc Anh, et al.
Veröffentlicht: (2025)
von: Vu, Duc Anh, et al.
Veröffentlicht: (2025)
Bias Evaluation and Mitigation in Retrieval-Augmented Medical Question-Answering Systems
von: Ji, Yuelyu, et al.
Veröffentlicht: (2025)
von: Ji, Yuelyu, et al.
Veröffentlicht: (2025)
Positional Bias in Binary Question Answering: How Uncertainty Shapes Model Preferences
von: Labruna, Tiziano, et al.
Veröffentlicht: (2025)
von: Labruna, Tiziano, et al.
Veröffentlicht: (2025)
Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language Models
von: Loginova, Olga, et al.
Veröffentlicht: (2024)
von: Loginova, Olga, et al.
Veröffentlicht: (2024)
Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions
von: Chen, Hanjie, et al.
Veröffentlicht: (2024)
von: Chen, Hanjie, et al.
Veröffentlicht: (2024)
BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models
von: Narnaware, Vishal, et al.
Veröffentlicht: (2025)
von: Narnaware, Vishal, et al.
Veröffentlicht: (2025)
KatotohananQA: Evaluating Truthfulness of Large Language Models in Filipino
von: Nery, Lorenzo Alfred, et al.
Veröffentlicht: (2025)
von: Nery, Lorenzo Alfred, et al.
Veröffentlicht: (2025)
Probability of Differentiation Reveals Brittleness of Homogeneity Bias in GPT-4
von: Lee, Messi H. J., et al.
Veröffentlicht: (2024)
von: Lee, Messi H. J., et al.
Veröffentlicht: (2024)
FREB-TQA: A Fine-Grained Robustness Evaluation Benchmark for Table Question Answering
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
McBE: A Multi-task Chinese Bias Evaluation Benchmark for Large Language Models
von: Lan, Tian, et al.
Veröffentlicht: (2025)
von: Lan, Tian, et al.
Veröffentlicht: (2025)
Large Language Models Portray Socially Subordinate Groups as More Homogeneous, Consistent with a Bias Observed in Humans
von: Lee, Messi H. J., et al.
Veröffentlicht: (2024)
von: Lee, Messi H. J., et al.
Veröffentlicht: (2024)
IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages
von: Hanif, Ikhlasul Akmal, et al.
Veröffentlicht: (2026)
von: Hanif, Ikhlasul Akmal, et al.
Veröffentlicht: (2026)
Measuring Agreeableness Bias in Multimodal Models
von: Lim, Jaehyuk, et al.
Veröffentlicht: (2024)
von: Lim, Jaehyuk, et al.
Veröffentlicht: (2024)
Gender Bias in Emotion Recognition by Large Language Models
von: Herbert, Maureen, et al.
Veröffentlicht: (2025)
von: Herbert, Maureen, et al.
Veröffentlicht: (2025)
Social Bias Probing: Fairness Benchmarking for Language Models
von: Manerba, Marta Marchiori, et al.
Veröffentlicht: (2023)
von: Manerba, Marta Marchiori, et al.
Veröffentlicht: (2023)
CLIMB: A Benchmark of Clinical Bias in Large Language Models
von: Zhang, Yubo, et al.
Veröffentlicht: (2024)
von: Zhang, Yubo, et al.
Veröffentlicht: (2024)
Bias in, Bias out: Annotation Bias in Multilingual Large Language Models
von: Cui, Xia, et al.
Veröffentlicht: (2025)
von: Cui, Xia, et al.
Veröffentlicht: (2025)
A Survey of Large Language Model Agents for Question Answering
von: Yue, Murong
Veröffentlicht: (2025)
von: Yue, Murong
Veröffentlicht: (2025)
Evaluating Gender Bias in Large Language Models
von: Döll, Michael, et al.
Veröffentlicht: (2024)
von: Döll, Michael, et al.
Veröffentlicht: (2024)
Mitigating the Bias of Large Language Model Evaluation
von: Zhou, Hongli, et al.
Veröffentlicht: (2024)
von: Zhou, Hongli, et al.
Veröffentlicht: (2024)
FLEX: A Benchmark for Evaluating Robustness of Fairness in Large Language Models
von: Jung, Dahyun, et al.
Veröffentlicht: (2025)
von: Jung, Dahyun, et al.
Veröffentlicht: (2025)
Obscured but Not Erased: Evaluating Nationality Bias in LLMs via Name-Based Bias Benchmarks
von: Pelosio, Giulio, et al.
Veröffentlicht: (2025)
von: Pelosio, Giulio, et al.
Veröffentlicht: (2025)
Argument-Based Comparative Question Answering Evaluation Benchmark
von: Nikishina, Irina, et al.
Veröffentlicht: (2025)
von: Nikishina, Irina, et al.
Veröffentlicht: (2025)
PAT-Questions: A Self-Updating Benchmark for Present-Anchored Temporal Question-Answering
von: Meem, Jannat Ara, et al.
Veröffentlicht: (2024)
von: Meem, Jannat Ara, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Bias Attribution in Filipino Language Models: Extending a Bias Interpretability Metric for Application on Agglutinative Languages
von: Gamboa, Lance Calvin Lim, et al.
Veröffentlicht: (2025) -
Filipino Benchmarks for Measuring Sexist and Homophobic Bias in Multilingual Language Models from Southeast Asia
von: Gamboa, Lance Calvin Lim, et al.
Veröffentlicht: (2024) -
Social Bias in Multilingual Language Models: A Survey
von: Gamboa, Lance Calvin Lim, et al.
Veröffentlicht: (2025) -
KoBBQ: Korean Bias Benchmark for Question Answering
von: Jin, Jiho, et al.
Veröffentlicht: (2023) -
A Novel Interpretability Metric for Explaining Bias in Language Models: Applications on Multilingual Models from Southeast Asia
von: Gamboa, Lance Calvin Lim, et al.
Veröffentlicht: (2024)