Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yang, Chu, Chenhui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Humans or LLMs as the Judge? A Study on Judgement Biases
by: Chen, Guiming Hardy, et al.
Published: (2024)
by: Chen, Guiming Hardy, et al.
Published: (2024)
LLMs are Biased Evaluators But Not Biased for Retrieval Augmented Generation
by: Chen, Yen-Shan, et al.
Published: (2024)
by: Chen, Yen-Shan, et al.
Published: (2024)
BanStereoSet: A Dataset to Measure Stereotypical Social Biases in LLMs for Bangla
by: Kamruzzaman, Mahammed, et al.
Published: (2024)
by: Kamruzzaman, Mahammed, et al.
Published: (2024)
"Not in My Backyard": LLMs Uncover Online and Offline Social Biases Against Homelessness
by: Karr Jr., Jonathan A., et al.
Published: (2025)
by: Karr Jr., Jonathan A., et al.
Published: (2025)
Identifying High-Confidence Social Biases in LLMs for Trustworthy Conversational Tutoring Agents
by: Alvarez, Aitor Arronte, et al.
Published: (2026)
by: Alvarez, Aitor Arronte, et al.
Published: (2026)
Breaking Bias, Building Bridges: Evaluation and Mitigation of Social Biases in LLMs via Contact Hypothesis
by: Raj, Chahat, et al.
Published: (2024)
by: Raj, Chahat, et al.
Published: (2024)
Robust Evaluation Measures for Evaluating Social Biases in Masked Language Models
by: Liu, Yang
Published: (2024)
by: Liu, Yang
Published: (2024)
LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation
by: Yang, Gao, et al.
Published: (2025)
by: Yang, Gao, et al.
Published: (2025)
The Biased Oracle: Assessing LLMs' Understandability and Empathy in Medical Diagnoses
by: Yao, Jianzhou, et al.
Published: (2025)
by: Yao, Jianzhou, et al.
Published: (2025)
MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning
by: Ghosh, Rajarshi, et al.
Published: (2025)
by: Ghosh, Rajarshi, et al.
Published: (2025)
GenderBench: Evaluation Suite for Gender Biases in LLMs
by: Pikuliak, Matúš
Published: (2025)
by: Pikuliak, Matúš
Published: (2025)
Analyzing Dialectical Biases in LLMs for Knowledge and Reasoning Benchmarks
by: Pan, Eileen, et al.
Published: (2025)
by: Pan, Eileen, et al.
Published: (2025)
Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages
by: Naous, Tarek, et al.
Published: (2025)
by: Naous, Tarek, et al.
Published: (2025)
A Comprehensive Evaluation of Cognitive Biases in LLMs
by: Malberg, Simon, et al.
Published: (2024)
by: Malberg, Simon, et al.
Published: (2024)
Demo: Statistically Significant Results On Biases and Errors of LLMs Do Not Guarantee Generalizable Results
by: Liu, Jonathan, et al.
Published: (2025)
by: Liu, Jonathan, et al.
Published: (2025)
White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMs
by: Wan, Yixin, et al.
Published: (2024)
by: Wan, Yixin, et al.
Published: (2024)
Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned Strategies
by: Song, Zirui, et al.
Published: (2025)
by: Song, Zirui, et al.
Published: (2025)
Evaluating the Effect of Retrieval Augmentation on Social Biases
by: Zhang, Tianhui, et al.
Published: (2025)
by: Zhang, Tianhui, et al.
Published: (2025)
Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs
by: Yang, Xikang, et al.
Published: (2025)
by: Yang, Xikang, et al.
Published: (2025)
Do Biased Models Have Biased Thoughts?
by: Rajwal, Swati, et al.
Published: (2025)
by: Rajwal, Swati, et al.
Published: (2025)
To Lie or Not to Lie? Investigating The Biased Spread of Global Lies by LLMs
by: Khan, Zohaib, et al.
Published: (2026)
by: Khan, Zohaib, et al.
Published: (2026)
LLMs are Biased Teachers: Evaluating LLM Bias in Personalized Education
by: Weissburg, Iain, et al.
Published: (2024)
by: Weissburg, Iain, et al.
Published: (2024)
Silenced Biases: The Dark Side LLMs Learned to Refuse
by: Himelstein, Rom, et al.
Published: (2025)
by: Himelstein, Rom, et al.
Published: (2025)
Unveiling Divergent Inductive Biases of LLMs on Temporal Data
by: Kishore, Sindhu, et al.
Published: (2024)
by: Kishore, Sindhu, et al.
Published: (2024)
LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMs
by: Long, Do Xuan, et al.
Published: (2024)
by: Long, Do Xuan, et al.
Published: (2024)
Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases
by: Huang, Hui, et al.
Published: (2026)
by: Huang, Hui, et al.
Published: (2026)
Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs
by: de Landa, Joseba Fernandez, et al.
Published: (2026)
by: de Landa, Joseba Fernandez, et al.
Published: (2026)
Generative Language Models Exhibit Social Identity Biases
by: Hu, Tiancheng, et al.
Published: (2023)
by: Hu, Tiancheng, et al.
Published: (2023)
What an Elegant Bridge: Multilingual LLMs are Biased Similarly in Different Languages
by: Mihaylov, Viktor, et al.
Published: (2024)
by: Mihaylov, Viktor, et al.
Published: (2024)
LLM economicus? Mapping the Behavioral Biases of LLMs via Utility Theory
by: Ross, Jillian, et al.
Published: (2024)
by: Ross, Jillian, et al.
Published: (2024)
Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
by: Gupta, Shashank, et al.
Published: (2023)
by: Gupta, Shashank, et al.
Published: (2023)
Robust Pronoun Fidelity with English LLMs: Are they Reasoning, Repeating, or Just Biased?
by: Gautam, Vagrant, et al.
Published: (2024)
by: Gautam, Vagrant, et al.
Published: (2024)
Split and Merge: Aligning Position Biases in LLM-based Evaluators
by: Li, Zongjie, et al.
Published: (2023)
by: Li, Zongjie, et al.
Published: (2023)
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
by: Ye, Jiayi, et al.
Published: (2024)
by: Ye, Jiayi, et al.
Published: (2024)
Evaluating and Aligning CodeLLMs on Human Preference
by: Yang, Jian, et al.
Published: (2024)
by: Yang, Jian, et al.
Published: (2024)
The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Pre-trained Language Models
by: Liu, Yan, et al.
Published: (2024)
by: Liu, Yan, et al.
Published: (2024)
Mitigating Social Biases in Language Models through Unlearning
by: Dige, Omkar, et al.
Published: (2024)
by: Dige, Omkar, et al.
Published: (2024)
LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
by: Hu, Xuhao, et al.
Published: (2025)
by: Hu, Xuhao, et al.
Published: (2025)
Location Not Found: Exposing Implicit Local and Global Biases in Multilingual LLMs
by: Mor-Lan, Guy, et al.
Published: (2026)
by: Mor-Lan, Guy, et al.
Published: (2026)
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
Similar Items
-
Humans or LLMs as the Judge? A Study on Judgement Biases
by: Chen, Guiming Hardy, et al.
Published: (2024) -
LLMs are Biased Evaluators But Not Biased for Retrieval Augmented Generation
by: Chen, Yen-Shan, et al.
Published: (2024) -
BanStereoSet: A Dataset to Measure Stereotypical Social Biases in LLMs for Bangla
by: Kamruzzaman, Mahammed, et al.
Published: (2024) -
"Not in My Backyard": LLMs Uncover Online and Offline Social Biases Against Homelessness
by: Karr Jr., Jonathan A., et al.
Published: (2025) -
Identifying High-Confidence Social Biases in LLMs for Trustworthy Conversational Tutoring Agents
by: Alvarez, Aitor Arronte, et al.
Published: (2026)