Addressing Stereotypes in Large Language Models: A Critical Examination and Mitigation
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Kazi, Fatima |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Taxonomy of Stereotype Content in Large Language Models
von: Nicolas, Gandalf, et al.
Veröffentlicht: (2024)
von: Nicolas, Gandalf, et al.
Veröffentlicht: (2024)
Self-Debiasing Large Language Models: Zero-Shot Recognition and Reduction of Stereotypes
von: Gallegos, Isabel O., et al.
Veröffentlicht: (2024)
von: Gallegos, Isabel O., et al.
Veröffentlicht: (2024)
Biased or Flawed? Mitigating Stereotypes in Generative Language Models by Addressing Task-Specific Flaws
von: Jha, Akshita, et al.
Veröffentlicht: (2024)
von: Jha, Akshita, et al.
Veröffentlicht: (2024)
BiasEdit: Debiasing Stereotyped Language Models via Model Editing
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
A Comprehensive Study of Implicit and Explicit Biases in Large Language Models
von: Kazi, Fatima, et al.
Veröffentlicht: (2025)
von: Kazi, Fatima, et al.
Veröffentlicht: (2025)
Intrinsic Meets Extrinsic Fairness: Assessing the Downstream Impact of Bias Mitigation in Large Language Models
von: Arzaghi', 'Mina, et al.
Veröffentlicht: (2025)
von: Arzaghi', 'Mina, et al.
Veröffentlicht: (2025)
Multilingual Text-to-Image Generation Magnifies Gender Stereotypes and Prompt Engineering May Not Help You
von: Friedrich, Felix, et al.
Veröffentlicht: (2024)
von: Friedrich, Felix, et al.
Veröffentlicht: (2024)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
Addressing Both Statistical and Causal Gender Fairness in NLP Models
von: Chen, Hannah, et al.
Veröffentlicht: (2024)
von: Chen, Hannah, et al.
Veröffentlicht: (2024)
On The Role of Reasoning in the Identification of Subtle Stereotypes in Natural Language
von: Tian, Jacob-Junqi, et al.
Veröffentlicht: (2023)
von: Tian, Jacob-Junqi, et al.
Veröffentlicht: (2023)
Towards Modeling Learner Performance with Large Language Models
von: Neshaei, Seyed Parsa, et al.
Veröffentlicht: (2024)
von: Neshaei, Seyed Parsa, et al.
Veröffentlicht: (2024)
Harnessing Large Language Models for Disaster Management: A Survey
von: Lei, Zhenyu, et al.
Veröffentlicht: (2025)
von: Lei, Zhenyu, et al.
Veröffentlicht: (2025)
Protected group bias and stereotypes in Large Language Models
von: Kotek, Hadas, et al.
Veröffentlicht: (2024)
von: Kotek, Hadas, et al.
Veröffentlicht: (2024)
Understanding Intrinsic Socioeconomic Biases in Large Language Models
von: Arzaghi, Mina, et al.
Veröffentlicht: (2024)
von: Arzaghi, Mina, et al.
Veröffentlicht: (2024)
A Toolbox for Surfacing Health Equity Harms and Biases in Large Language Models
von: Pfohl, Stephen R., et al.
Veröffentlicht: (2024)
von: Pfohl, Stephen R., et al.
Veröffentlicht: (2024)
Limits to Predicting Online Speech Using Large Language Models
von: Remeli, Mina, et al.
Veröffentlicht: (2024)
von: Remeli, Mina, et al.
Veröffentlicht: (2024)
Revealing Fine-Grained Values and Opinions in Large Language Models
von: Wright, Dustin, et al.
Veröffentlicht: (2024)
von: Wright, Dustin, et al.
Veröffentlicht: (2024)
ELMES: An Automated Framework for Evaluating Large Language Models in Educational Scenarios
von: Wei, Shou'ang, et al.
Veröffentlicht: (2025)
von: Wei, Shou'ang, et al.
Veröffentlicht: (2025)
Generalization in Healthcare AI: Evaluation of a Clinical Large Language Model
von: Rahman, Salman, et al.
Veröffentlicht: (2024)
von: Rahman, Salman, et al.
Veröffentlicht: (2024)
A Detailed Factor Analysis for the Political Compass Test: Navigating Ideologies of Large Language Models
von: Kamal, Sadia, et al.
Veröffentlicht: (2025)
von: Kamal, Sadia, et al.
Veröffentlicht: (2025)
LIBRA: Measuring Bias of Large Language Model from a Local Context
von: Pang, Bo, et al.
Veröffentlicht: (2025)
von: Pang, Bo, et al.
Veröffentlicht: (2025)
Exploring the Potential of the Large Language Models (LLMs) in Identifying Misleading News Headlines
von: Rony, Md Main Uddin, et al.
Veröffentlicht: (2024)
von: Rony, Md Main Uddin, et al.
Veröffentlicht: (2024)
Cultural Alignment in Large Language Models: An Explanatory Analysis Based on Hofstede's Cultural Dimensions
von: Masoud, Reem I., et al.
Veröffentlicht: (2023)
von: Masoud, Reem I., et al.
Veröffentlicht: (2023)
What Large Language Models Do Not Talk About: An Empirical Study of Moderation and Censorship Practices
von: Noels, Sander, et al.
Veröffentlicht: (2025)
von: Noels, Sander, et al.
Veröffentlicht: (2025)
Estimating Item Difficulty Using Large Language Models and Tree-Based Machine Learning Algorithms
von: Razavi, Pooya, et al.
Veröffentlicht: (2025)
von: Razavi, Pooya, et al.
Veröffentlicht: (2025)
Augmenting Human-Annotated Training Data with Large Language Model Generation and Distillation in Open-Response Assessment
von: Borchers, Conrad, et al.
Veröffentlicht: (2025)
von: Borchers, Conrad, et al.
Veröffentlicht: (2025)
Automating Governing Knowledge Commons and Contextual Integrity (GKC-CI) Privacy Policy Annotations with Large Language Models
von: Chanenson, Jake, et al.
Veröffentlicht: (2023)
von: Chanenson, Jake, et al.
Veröffentlicht: (2023)
Bias and Fairness in Large Language Models: A Survey
von: Gallegos, Isabel O., et al.
Veröffentlicht: (2023)
von: Gallegos, Isabel O., et al.
Veröffentlicht: (2023)
DSO: Direct Steering Optimization for Bias Mitigation
von: Paes, Lucas Monteiro, et al.
Veröffentlicht: (2025)
von: Paes, Lucas Monteiro, et al.
Veröffentlicht: (2025)
The Moral Gap of Large Language Models
von: Skorski, Maciej, et al.
Veröffentlicht: (2025)
von: Skorski, Maciej, et al.
Veröffentlicht: (2025)
Correlated Errors in Large Language Models
von: Kim, Elliot, et al.
Veröffentlicht: (2025)
von: Kim, Elliot, et al.
Veröffentlicht: (2025)
Hypothesis Generation with Large Language Models
von: Zhou, Yangqiaoyu, et al.
Veröffentlicht: (2024)
von: Zhou, Yangqiaoyu, et al.
Veröffentlicht: (2024)
Large Language Models are Geographically Biased
von: Manvi, Rohin, et al.
Veröffentlicht: (2024)
von: Manvi, Rohin, et al.
Veröffentlicht: (2024)
Exploring the Linear Subspace Hypothesis in Gender Bias Mitigation
von: Vargas, Francisco, et al.
Veröffentlicht: (2020)
von: Vargas, Francisco, et al.
Veröffentlicht: (2020)
Fair Representation in Parliamentary Summaries: Measuring and Mitigating Inclusion Bias
von: Cunningham, Eoghan, et al.
Veröffentlicht: (2025)
von: Cunningham, Eoghan, et al.
Veröffentlicht: (2025)
Psychological Counseling Ability of Large Language Models
von: Peng, Fangyu, et al.
Veröffentlicht: (2025)
von: Peng, Fangyu, et al.
Veröffentlicht: (2025)
Assessing Large Language Models on Climate Information
von: Bulian, Jannis, et al.
Veröffentlicht: (2023)
von: Bulian, Jannis, et al.
Veröffentlicht: (2023)
A Moral Imperative: The Need for Continual Superalignment of Large Language Models
von: Puthumanaillam, Gokul, et al.
Veröffentlicht: (2024)
von: Puthumanaillam, Gokul, et al.
Veröffentlicht: (2024)
Leveraging Prototypical Representations for Mitigating Social Bias without Demographic Information
von: Iskander, Shadi, et al.
Veröffentlicht: (2024)
von: Iskander, Shadi, et al.
Veröffentlicht: (2024)
Analyzing the Safety of Japanese Large Language Models in Stereotype-Triggering Prompts
von: Nakanishi, Akito, et al.
Veröffentlicht: (2025)
von: Nakanishi, Akito, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Taxonomy of Stereotype Content in Large Language Models
von: Nicolas, Gandalf, et al.
Veröffentlicht: (2024) -
Self-Debiasing Large Language Models: Zero-Shot Recognition and Reduction of Stereotypes
von: Gallegos, Isabel O., et al.
Veröffentlicht: (2024) -
Biased or Flawed? Mitigating Stereotypes in Generative Language Models by Addressing Task-Specific Flaws
von: Jha, Akshita, et al.
Veröffentlicht: (2024) -
BiasEdit: Debiasing Stereotyped Language Models via Model Editing
von: Xu, Xin, et al.
Veröffentlicht: (2025) -
A Comprehensive Study of Implicit and Explicit Biases in Large Language Models
von: Kazi, Fatima, et al.
Veröffentlicht: (2025)