Social Bias Benchmark for Generation: A Comparison of Generation and QA-Based Evaluations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jin, Jiho, Kang, Woosung, Myung, Junho, Oh, Alice |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
von: Jin, Jiho, et al.
Veröffentlicht: (2026)
von: Jin, Jiho, et al.
Veröffentlicht: (2026)
KoBBQ: Korean Bias Benchmark for Question Answering
von: Jin, Jiho, et al.
Veröffentlicht: (2023)
von: Jin, Jiho, et al.
Veröffentlicht: (2023)
Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis
von: Lee, Nayeon, et al.
Veröffentlicht: (2023)
von: Lee, Nayeon, et al.
Veröffentlicht: (2023)
Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models
von: Yu, Haeun, et al.
Veröffentlicht: (2025)
von: Yu, Haeun, et al.
Veröffentlicht: (2025)
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language
von: Song, Seyoung, et al.
Veröffentlicht: (2025)
von: Song, Seyoung, et al.
Veröffentlicht: (2025)
Code-Switching In-Context Learning for Cross-Lingual Transfer of Large Language Models
von: Yoo, Haneul, et al.
Veröffentlicht: (2025)
von: Yoo, Haneul, et al.
Veröffentlicht: (2025)
The Generative AI Paradox on Evaluation: What It Can Solve, It May Not Evaluate
von: Oh, Juhyun, et al.
Veröffentlicht: (2024)
von: Oh, Juhyun, et al.
Veröffentlicht: (2024)
GRADE: Generating multi-hop QA and fine-gRAined Difficulty matrix for RAG Evaluation
von: Lee, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Lee, Jeongsoo, et al.
Veröffentlicht: (2025)
PakBBQ: A Culturally Adapted Bias Benchmark for QA
von: Hashmat, Abdullah, et al.
Veröffentlicht: (2025)
von: Hashmat, Abdullah, et al.
Veröffentlicht: (2025)
FINEST: Improving LLM Responses to Sensitive Topics Through Fine-Grained Evaluation
von: Oh, Juhyun, et al.
Veröffentlicht: (2026)
von: Oh, Juhyun, et al.
Veröffentlicht: (2026)
Translating Hanja Historical Documents to Contemporary Korean and English
von: Son, Juhee, et al.
Veröffentlicht: (2022)
von: Son, Juhee, et al.
Veröffentlicht: (2022)
PapersPlease: A Benchmark for Evaluating Motivational Values of Large Language Models Based on ERG Theory
von: Myung, Junho, et al.
Veröffentlicht: (2025)
von: Myung, Junho, et al.
Veröffentlicht: (2025)
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
von: Demchak, Nathaniel, et al.
Veröffentlicht: (2024)
von: Demchak, Nathaniel, et al.
Veröffentlicht: (2024)
M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
von: Li, Chuhan, et al.
Veröffentlicht: (2024)
von: Li, Chuhan, et al.
Veröffentlicht: (2024)
AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation
von: Huang, Tiancheng, et al.
Veröffentlicht: (2025)
von: Huang, Tiancheng, et al.
Veröffentlicht: (2025)
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
von: Shin, Jisu, et al.
Veröffentlicht: (2025)
von: Shin, Jisu, et al.
Veröffentlicht: (2025)
Generating Leakage-Free Benchmarks for Robust RAG Evaluation
von: Liu, Jiayi, et al.
Veröffentlicht: (2026)
von: Liu, Jiayi, et al.
Veröffentlicht: (2026)
Suvach -- Generated Hindi QA benchmark
von: Narayanan, Vaishak, et al.
Veröffentlicht: (2024)
von: Narayanan, Vaishak, et al.
Veröffentlicht: (2024)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
von: Shin, Jisu, et al.
Veröffentlicht: (2025)
von: Shin, Jisu, et al.
Veröffentlicht: (2025)
Arena-Lite: Efficient and Reliable Large Language Model Evaluation via Tournament-Based Direct Comparisons
von: Son, Seonil, et al.
Veröffentlicht: (2024)
von: Son, Seonil, et al.
Veröffentlicht: (2024)
MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs
von: Wei, Jianhui, et al.
Veröffentlicht: (2025)
von: Wei, Jianhui, et al.
Veröffentlicht: (2025)
DesignQA: A Multimodal Benchmark for Evaluating Large Language Models' Understanding of Engineering Documentation
von: Doris, Anna C., et al.
Veröffentlicht: (2024)
von: Doris, Anna C., et al.
Veröffentlicht: (2024)
KGQuest: Template-Driven QA Generation from Knowledge Graphs with LLM-Based Refinement
von: Nayab, Sania, et al.
Veröffentlicht: (2025)
von: Nayab, Sania, et al.
Veröffentlicht: (2025)
LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA
von: Bonomo, Tommaso, et al.
Veröffentlicht: (2025)
von: Bonomo, Tommaso, et al.
Veröffentlicht: (2025)
Social Bias in Popular Question-Answering Benchmarks
von: Kraft, Angelie, et al.
Veröffentlicht: (2025)
von: Kraft, Angelie, et al.
Veröffentlicht: (2025)
Threads of Subtlety: Detecting Machine-Generated Texts Through Discourse Motifs
von: Kim, Zae Myung, et al.
Veröffentlicht: (2024)
von: Kim, Zae Myung, et al.
Veröffentlicht: (2024)
R2-KG: General-Purpose Dual-Agent Framework for Reliable Reasoning on Knowledge Graphs
von: Jo, Sumin, et al.
Veröffentlicht: (2025)
von: Jo, Sumin, et al.
Veröffentlicht: (2025)
Evaluating Monolingual and Multilingual Large Language Models for Greek Question Answering: The DemosQA Benchmark
von: Mastrokostas, Charalampos, et al.
Veröffentlicht: (2026)
von: Mastrokostas, Charalampos, et al.
Veröffentlicht: (2026)
IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages
von: Hanif, Ikhlasul Akmal, et al.
Veröffentlicht: (2026)
von: Hanif, Ikhlasul Akmal, et al.
Veröffentlicht: (2026)
CEval: A Benchmark for Evaluating Counterfactual Text Generation
von: Nguyen, Van Bach, et al.
Veröffentlicht: (2024)
von: Nguyen, Van Bach, et al.
Veröffentlicht: (2024)
Benchmark Test-Time Scaling of General LLM Agents
von: Li, Xiaochuan, et al.
Veröffentlicht: (2026)
von: Li, Xiaochuan, et al.
Veröffentlicht: (2026)
Does Reasoning Introduce Bias? A Study of Social Bias Evaluation and Mitigation in LLM Reasoning
von: Wu, Xuyang, et al.
Veröffentlicht: (2025)
von: Wu, Xuyang, et al.
Veröffentlicht: (2025)
Generating Benchmarks for Factuality Evaluation of Language Models
von: Muhlgay, Dor, et al.
Veröffentlicht: (2023)
von: Muhlgay, Dor, et al.
Veröffentlicht: (2023)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
von: Kim, Eunsu, et al.
Veröffentlicht: (2025)
von: Kim, Eunsu, et al.
Veröffentlicht: (2025)
Long-Span Question-Answering: Automatic Question Generation and QA-System Ranking via Side-by-Side Evaluation
von: Bohnet, Bernd, et al.
Veröffentlicht: (2024)
von: Bohnet, Bernd, et al.
Veröffentlicht: (2024)
Benchmarking Retrieval-Augmented Generation for Medicine
von: Xiong, Guangzhi, et al.
Veröffentlicht: (2024)
von: Xiong, Guangzhi, et al.
Veröffentlicht: (2024)
SensorQA: A Question Answering Benchmark for Daily-Life Monitoring
von: Reichman, Benjamin, et al.
Veröffentlicht: (2025)
von: Reichman, Benjamin, et al.
Veröffentlicht: (2025)
Developing A Framework to Support Human Evaluation of Bias in Generated Free Response Text
von: Healey, Jennifer, et al.
Veröffentlicht: (2025)
von: Healey, Jennifer, et al.
Veröffentlicht: (2025)
MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering
von: Islamaj, Rezarta, et al.
Veröffentlicht: (2026)
von: Islamaj, Rezarta, et al.
Veröffentlicht: (2026)
Man Made Language Models? Evaluating LLMs' Perpetuation of Masculine Generics Bias
von: Doyen, Enzo, et al.
Veröffentlicht: (2025)
von: Doyen, Enzo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
von: Jin, Jiho, et al.
Veröffentlicht: (2026) -
KoBBQ: Korean Bias Benchmark for Question Answering
von: Jin, Jiho, et al.
Veröffentlicht: (2023) -
Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis
von: Lee, Nayeon, et al.
Veröffentlicht: (2023) -
Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models
von: Yu, Haeun, et al.
Veröffentlicht: (2025) -
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language
von: Song, Seyoung, et al.
Veröffentlicht: (2025)