How Are LLMs Mitigating Stereotyping Harms? Learning from Search Engine Studies
Fuente:
arXiv
Guardado en:
| Autores principales: | Leidinger, Alina, Rogers, Richard |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Are LLMs classical or nonmonotonic reasoners? Lessons from generics
por: Leidinger, Alina, et al.
Publicado: (2024)
por: Leidinger, Alina, et al.
Publicado: (2024)
How far can bias go? Tracing bias from pretraining data to alignment
por: Thaler, Marion, et al.
Publicado: (2024)
por: Thaler, Marion, et al.
Publicado: (2024)
Beating Harmful Stereotypes Through Facts: RAG-based Counter-speech Generation
por: Damo, Greta, et al.
Publicado: (2025)
por: Damo, Greta, et al.
Publicado: (2025)
LLMs Reproduce Stereotypes of Sexual and Gender Minorities
por: Ostrow, Ruby, et al.
Publicado: (2025)
por: Ostrow, Ruby, et al.
Publicado: (2025)
REFINE-LM: Mitigating Language Model Stereotypes via Reinforcement Learning
por: Qureshi, Rameez, et al.
Publicado: (2024)
por: Qureshi, Rameez, et al.
Publicado: (2024)
Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
por: Jin, Bowen, et al.
Publicado: (2025)
por: Jin, Bowen, et al.
Publicado: (2025)
Probing Gender Bias in Multilingual LLMs: A Case Study of Stereotypes in Persian
por: Kalhor, Ghazal, et al.
Publicado: (2025)
por: Kalhor, Ghazal, et al.
Publicado: (2025)
Do Multilingual Large Language Models Mitigate Stereotype Bias?
por: Nie, Shangrui, et al.
Publicado: (2024)
por: Nie, Shangrui, et al.
Publicado: (2024)
DART: Mitigating Harm Drift in Difference-Aware LLMs via Distill-Audit-Repair Training
por: Pan, Ziwen, et al.
Publicado: (2026)
por: Pan, Ziwen, et al.
Publicado: (2026)
`For Argument's Sake, Show Me How to Harm Myself!': Jailbreaking LLMs in Suicide and Self-Harm Contexts
por: Schoene, Annika M, et al.
Publicado: (2025)
por: Schoene, Annika M, et al.
Publicado: (2025)
CIVICS: Building a Dataset for Examining Culturally-Informed Values in Large Language Models
por: Pistilli, Giada, et al.
Publicado: (2024)
por: Pistilli, Giada, et al.
Publicado: (2024)
Challenging Negative Gender Stereotypes: A Study on the Effectiveness of Automated Counter-Stereotypes
por: Nejadgholi, Isar, et al.
Publicado: (2024)
por: Nejadgholi, Isar, et al.
Publicado: (2024)
Can We Locate and Prevent Stereotypes in LLMs?
por: D'Souza, Alex
Publicado: (2026)
por: D'Souza, Alex
Publicado: (2026)
Profiling Bias in LLMs: Stereotype Dimensions in Contextual Word Embeddings
por: Schuster, Carolin M., et al.
Publicado: (2024)
por: Schuster, Carolin M., et al.
Publicado: (2024)
Can Editing LLMs Inject Harm?
por: Chen, Canyu, et al.
Publicado: (2024)
por: Chen, Canyu, et al.
Publicado: (2024)
LLMs Encode Harmfulness and Refusal Separately
por: Zhao, Jiachen, et al.
Publicado: (2025)
por: Zhao, Jiachen, et al.
Publicado: (2025)
Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
por: Atil, Berk, et al.
Publicado: (2025)
por: Atil, Berk, et al.
Publicado: (2025)
Mitigating Harmful Erraticism in LLMs Through Dialectical Behavior Therapy Based De-Escalation Strategies
por: Rangarajan, Pooja, et al.
Publicado: (2025)
por: Rangarajan, Pooja, et al.
Publicado: (2025)
Undesirable Biases in NLP: Addressing Challenges of Measurement
por: van der Wal, Oskar, et al.
Publicado: (2022)
por: van der Wal, Oskar, et al.
Publicado: (2022)
MBBQ: A Dataset for Cross-Lingual Comparison of Stereotypes in Generative LLMs
por: Neplenbroek, Vera, et al.
Publicado: (2024)
por: Neplenbroek, Vera, et al.
Publicado: (2024)
JUBAKU: An Adversarial Benchmark for Exposing Culturally Grounded Stereotypes in Japanese LLMs
por: Shiotani, Taihei, et al.
Publicado: (2026)
por: Shiotani, Taihei, et al.
Publicado: (2026)
Addressing Stereotypes in Large Language Models: A Critical Examination and Mitigation
por: Kazi, Fatima
Publicado: (2025)
por: Kazi, Fatima
Publicado: (2025)
Are Stereotypes Leading LLMs' Zero-Shot Stance Detection ?
por: Dubreuil, Anthony, et al.
Publicado: (2025)
por: Dubreuil, Anthony, et al.
Publicado: (2025)
Ranking Manipulation for Conversational Search Engines
por: Pfrommer, Samuel, et al.
Publicado: (2024)
por: Pfrommer, Samuel, et al.
Publicado: (2024)
Reading Between the Prompts: How Stereotypes Shape LLM's Implicit Personalization
por: Neplenbroek, Vera, et al.
Publicado: (2025)
por: Neplenbroek, Vera, et al.
Publicado: (2025)
Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning
por: Feng, Weitao, et al.
Publicado: (2025)
por: Feng, Weitao, et al.
Publicado: (2025)
Surfacing Subtle Stereotypes: A Multilingual, Debate-Oriented Evaluation of Modern LLMs
por: Saeed, Muhammed, et al.
Publicado: (2025)
por: Saeed, Muhammed, et al.
Publicado: (2025)
Social and Political Framing in Search Engine Results
por: Poudel, Amrit, et al.
Publicado: (2025)
por: Poudel, Amrit, et al.
Publicado: (2025)
Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models
por: Dogra, Atharvan, et al.
Publicado: (2025)
por: Dogra, Atharvan, et al.
Publicado: (2025)
Do Prevalent Bias Metrics Capture Allocational Harms from LLMs?
por: Cyberey, Hannah, et al.
Publicado: (2024)
por: Cyberey, Hannah, et al.
Publicado: (2024)
BanStereoSet: A Dataset to Measure Stereotypical Social Biases in LLMs for Bangla
por: Kamruzzaman, Mahammed, et al.
Publicado: (2024)
por: Kamruzzaman, Mahammed, et al.
Publicado: (2024)
Quantifying Stereotypes in Language
por: Liu, Yang
Publicado: (2024)
por: Liu, Yang
Publicado: (2024)
Stereotype Detection in LLMs: A Multiclass, Explainable, and Benchmark-Driven Approach
por: Wu, Zekun, et al.
Publicado: (2024)
por: Wu, Zekun, et al.
Publicado: (2024)
Biased or Flawed? Mitigating Stereotypes in Generative Language Models by Addressing Task-Specific Flaws
por: Jha, Akshita, et al.
Publicado: (2024)
por: Jha, Akshita, et al.
Publicado: (2024)
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
por: Zhang, Chi, et al.
Publicado: (2025)
por: Zhang, Chi, et al.
Publicado: (2025)
Large Language Models as Search Engines: Societal Challenges
por: Sadeddine, Zacchary, et al.
Publicado: (2025)
por: Sadeddine, Zacchary, et al.
Publicado: (2025)
On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs
por: Ghorbanpour, Faeze, et al.
Publicado: (2025)
por: Ghorbanpour, Faeze, et al.
Publicado: (2025)
An Evaluation of LLMs for Detecting Harmful Computing Terms
por: Jacas, Joshua, et al.
Publicado: (2025)
por: Jacas, Joshua, et al.
Publicado: (2025)
Generative Engine Optimization: How to Dominate AI Search
por: Chen, Mahe, et al.
Publicado: (2025)
por: Chen, Mahe, et al.
Publicado: (2025)
StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs
por: Jeune, Pierre Le, et al.
Publicado: (2026)
por: Jeune, Pierre Le, et al.
Publicado: (2026)
Ejemplares similares
-
Are LLMs classical or nonmonotonic reasoners? Lessons from generics
por: Leidinger, Alina, et al.
Publicado: (2024) -
How far can bias go? Tracing bias from pretraining data to alignment
por: Thaler, Marion, et al.
Publicado: (2024) -
Beating Harmful Stereotypes Through Facts: RAG-based Counter-speech Generation
por: Damo, Greta, et al.
Publicado: (2025) -
LLMs Reproduce Stereotypes of Sexual and Gender Minorities
por: Ostrow, Ruby, et al.
Publicado: (2025) -
REFINE-LM: Mitigating Language Model Stereotypes via Reinforcement Learning
por: Qureshi, Rameez, et al.
Publicado: (2024)