Measuring Implicit Bias in Explicitly Unbiased Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bai, Xuechunzi, Wang, Angelina, Sucholutsky, Ilia, Griffiths, Thomas L. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Large Language Models Develop Novel Social Biases Through Adaptive Exploration
von: Wu, Addison J., et al.
Veröffentlicht: (2025)
von: Wu, Addison J., et al.
Veröffentlicht: (2025)
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race
von: Sun, Lihao, et al.
Veröffentlicht: (2025)
von: Sun, Lihao, et al.
Veröffentlicht: (2025)
Large Language Models Assume People are More Rational than We Really are
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
What is a Number, That a Large Language Model May Know It?
von: Marjieh, Raja, et al.
Veröffentlicht: (2025)
von: Marjieh, Raja, et al.
Veröffentlicht: (2025)
Covert Bias: The Severity of Social Views' Unalignment in Language Models Towards Implicit and Explicit Opinion
von: Aldayel, Abeer, et al.
Veröffentlicht: (2024)
von: Aldayel, Abeer, et al.
Veröffentlicht: (2024)
Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
Measuring Machine Learning Harms from Stereotypes Requires Understanding Who Is Harmed by Which Errors in What Ways
von: Wang, Angelina, et al.
Veröffentlicht: (2024)
von: Wang, Angelina, et al.
Veröffentlicht: (2024)
Inference-Time Reasoning Selectively Reduces Implicit Social Bias in Large Language Models
von: Apsel, Molly, et al.
Veröffentlicht: (2026)
von: Apsel, Molly, et al.
Veröffentlicht: (2026)
Levels of Analysis for Large Language Models
von: Ku, Alexander Y., et al.
Veröffentlicht: (2025)
von: Ku, Alexander Y., et al.
Veröffentlicht: (2025)
Leveraging Large Language Models to Measure Gender Representation Bias in Gendered Language Corpora
von: Derner, Erik, et al.
Veröffentlicht: (2024)
von: Derner, Erik, et al.
Veröffentlicht: (2024)
Analyzing the Roles of Language and Vision in Learning from Limited Data
von: Chen, Allison, et al.
Veröffentlicht: (2024)
von: Chen, Allison, et al.
Veröffentlicht: (2024)
Only a Little to the Left: A Theory-grounded Measure of Political Bias in Large Language Models
von: Faulborn, Mats, et al.
Veröffentlicht: (2025)
von: Faulborn, Mats, et al.
Veröffentlicht: (2025)
LIBRA: Measuring Bias of Large Language Model from a Local Context
von: Pang, Bo, et al.
Veröffentlicht: (2025)
von: Pang, Bo, et al.
Veröffentlicht: (2025)
Characterizing Selective Refusal Bias in Large Language Models
von: Khorramrouz, Adel, et al.
Veröffentlicht: (2025)
von: Khorramrouz, Adel, et al.
Veröffentlicht: (2025)
Gender Bias in Emotion Recognition by Large Language Models
von: Herbert, Maureen, et al.
Veröffentlicht: (2025)
von: Herbert, Maureen, et al.
Veröffentlicht: (2025)
The Algorithmic Unconscious: Structural Mechanisms and Implicit Biases in Large Language Models
von: Boisnard, Philippe
Veröffentlicht: (2026)
von: Boisnard, Philippe
Veröffentlicht: (2026)
Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language Models
von: Jhaveri, Ayush Rajesh, et al.
Veröffentlicht: (2026)
von: Jhaveri, Ayush Rajesh, et al.
Veröffentlicht: (2026)
VisBias: Measuring Explicit and Implicit Social Biases in Vision Language Models
von: Huang, Jen-tse, et al.
Veröffentlicht: (2025)
von: Huang, Jen-tse, et al.
Veröffentlicht: (2025)
Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
von: Wu, Addison J., et al.
Veröffentlicht: (2026)
von: Wu, Addison J., et al.
Veröffentlicht: (2026)
Cross-Language Bias Examination in Large Language Models
von: Liang, Yuxuan, et al.
Veröffentlicht: (2025)
von: Liang, Yuxuan, et al.
Veröffentlicht: (2025)
Measuring Large Language Models Capacity to Annotate Journalistic Sourcing
von: Vincent, Subramaniam, et al.
Veröffentlicht: (2024)
von: Vincent, Subramaniam, et al.
Veröffentlicht: (2024)
MALIBU Benchmark: Multi-Agent LLM Implicit Bias Uncovered
von: Mirza, Imran, et al.
Veröffentlicht: (2025)
von: Mirza, Imran, et al.
Veröffentlicht: (2025)
Ask LLMs Directly, "What shapes your bias?": Measuring Social Bias in Large Language Models
von: Shin, Jisu, et al.
Veröffentlicht: (2024)
von: Shin, Jisu, et al.
Veröffentlicht: (2024)
Towards Equitable AI: Detecting Bias in Using Large Language Models for Marketing
von: Yilmaz, Berk, et al.
Veröffentlicht: (2025)
von: Yilmaz, Berk, et al.
Veröffentlicht: (2025)
Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese
von: Lyu, Hanjia, et al.
Veröffentlicht: (2025)
von: Lyu, Hanjia, et al.
Veröffentlicht: (2025)
Towards Implicit Bias Detection and Mitigation in Multi-Agent LLM Interactions
von: Borah, Angana, et al.
Veröffentlicht: (2024)
von: Borah, Angana, et al.
Veröffentlicht: (2024)
Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs
von: Wang, Angelina, et al.
Veröffentlicht: (2025)
von: Wang, Angelina, et al.
Veröffentlicht: (2025)
CDEval: A Benchmark for Measuring the Cultural Dimensions of Large Language Models
von: Wang, Yuhang, et al.
Veröffentlicht: (2023)
von: Wang, Yuhang, et al.
Veröffentlicht: (2023)
Translate With Care: Addressing Gender Bias, Neutrality, and Reasoning in Large Language Model Translations
von: Zahraei, Pardis Sadat, et al.
Veröffentlicht: (2025)
von: Zahraei, Pardis Sadat, et al.
Veröffentlicht: (2025)
Explicit vs. Implicit: Investigating Social Bias in Large Language Models through Self-Reflection
von: Zhao, Yachao, et al.
Veröffentlicht: (2025)
von: Zhao, Yachao, et al.
Veröffentlicht: (2025)
Assessing Judging Bias in Large Reasoning Models: An Empirical Study
von: Wang, Qian, et al.
Veröffentlicht: (2025)
von: Wang, Qian, et al.
Veröffentlicht: (2025)
Divine LLaMAs: Bias, Stereotypes, Stigmatization, and Emotion Representation of Religion in Large Language Models
von: Plaza-del-Arco, Flor Miriam, et al.
Veröffentlicht: (2024)
von: Plaza-del-Arco, Flor Miriam, et al.
Veröffentlicht: (2024)
Down the Toxicity Rabbit Hole: A Novel Framework to Bias Audit Large Language Models
von: Dutta, Arka, et al.
Veröffentlicht: (2023)
von: Dutta, Arka, et al.
Veröffentlicht: (2023)
WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+ Bias in Large Language Models
von: Felkner, Virginia K., et al.
Veröffentlicht: (2023)
von: Felkner, Virginia K., et al.
Veröffentlicht: (2023)
DECASTE: Unveiling Caste Stereotypes in Large Language Models through Multi-Dimensional Bias Analysis
von: Vijayaraghavan, Prashanth, et al.
Veröffentlicht: (2025)
von: Vijayaraghavan, Prashanth, et al.
Veröffentlicht: (2025)
Gender Bias in Machine Translation and The Era of Large Language Models
von: Vanmassenhove, Eva
Veröffentlicht: (2024)
von: Vanmassenhove, Eva
Veröffentlicht: (2024)
AccessEval: Benchmarking Disability Bias in Large Language Models
von: Panda, Srikant, et al.
Veröffentlicht: (2025)
von: Panda, Srikant, et al.
Veröffentlicht: (2025)
Bias and Volatility: A Statistical Framework for Evaluating Large Language Model's Stereotypes and the Associated Generation Inconsistency
von: Liu, Yiran, et al.
Veröffentlicht: (2024)
von: Liu, Yiran, et al.
Veröffentlicht: (2024)
Dialect vs Demographics: Quantifying LLM Bias from Implicit Linguistic Signals vs. Explicit User Profiles
von: Haq, Irti, et al.
Veröffentlicht: (2026)
von: Haq, Irti, et al.
Veröffentlicht: (2026)
Agentic Society: Merging skeleton from real world and texture from Large Language Model
von: Bai, Yuqi, et al.
Veröffentlicht: (2024)
von: Bai, Yuqi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Large Language Models Develop Novel Social Biases Through Adaptive Exploration
von: Wu, Addison J., et al.
Veröffentlicht: (2025) -
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race
von: Sun, Lihao, et al.
Veröffentlicht: (2025) -
Large Language Models Assume People are More Rational than We Really are
von: Liu, Ryan, et al.
Veröffentlicht: (2024) -
What is a Number, That a Large Language Model May Know It?
von: Marjieh, Raja, et al.
Veröffentlicht: (2025) -
Covert Bias: The Severity of Social Views' Unalignment in Language Models Towards Implicit and Explicit Opinion
von: Aldayel, Abeer, et al.
Veröffentlicht: (2024)