First, Do No Harm (With LLMs): Mitigating Racial Bias via Agentic Workflows
Fuente:
arXiv
Saved in:
| Main Authors: | Xing, Sihao, Gouliev, Zaur |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Propaganda and Information Dissemination in the Russo-Ukrainian War: Natural Language Processing of Russian and Western Twitter Narratives
by: Gouliev, Zaur
Published: (2025)
by: Gouliev, Zaur
Published: (2025)
Uncovering Bias in Foundation Models: Impact, Testing, Harm, and Mitigation
by: Sun, Shuzhou, et al.
Published: (2025)
by: Sun, Shuzhou, et al.
Published: (2025)
Mitigating Gender Bias via Fostering Exploratory Thinking in LLMs
by: Wei, Kangda, et al.
Published: (2025)
by: Wei, Kangda, et al.
Published: (2025)
Evolution of AI in Education: Agentic Workflows
by: Kamalov, Firuz, et al.
Published: (2025)
by: Kamalov, Firuz, et al.
Published: (2025)
Echoes of AI Harms: A Human-LLM Synergistic Framework for Bias-Driven Harm Anticipation
by: Tantalaki, Nicoleta, et al.
Published: (2025)
by: Tantalaki, Nicoleta, et al.
Published: (2025)
Agentic Workflow for Education: Concepts and Applications
by: Jiang, Yuan-Hao, et al.
Published: (2025)
by: Jiang, Yuan-Hao, et al.
Published: (2025)
A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI
by: El-Sayed, Seliem, et al.
Published: (2024)
by: El-Sayed, Seliem, et al.
Published: (2024)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
by: Choi, Sooyung, et al.
Published: (2025)
by: Choi, Sooyung, et al.
Published: (2025)
From Bias Mitigation to Bias Negotiation: Governing Identity and Sociocultural Reasoning in Generative AI
by: Dunivin, Zackary Okun, et al.
Published: (2026)
by: Dunivin, Zackary Okun, et al.
Published: (2026)
Impacts of Racial Bias in Historical Training Data for News AI
by: Bhargava, Rahul, et al.
Published: (2025)
by: Bhargava, Rahul, et al.
Published: (2025)
Implicit Bias in LLMs for Transgender Populations
by: Hirsch, Micaela, et al.
Published: (2026)
by: Hirsch, Micaela, et al.
Published: (2026)
Measuring and Mitigating Bias in Code Generated by Large Language Models
by: Chen, Yuxi, et al.
Published: (2026)
by: Chen, Yuxi, et al.
Published: (2026)
Mitigating Gender Bias in Depression Detection via Counterfactual Inference
by: Hu, Mingxuan, et al.
Published: (2025)
by: Hu, Mingxuan, et al.
Published: (2025)
Adaptive Generation of Bias-Eliciting Questions for LLMs
by: Staab, Robin, et al.
Published: (2025)
by: Staab, Robin, et al.
Published: (2025)
Beyond Behaviorist Representational Harms: A Plan for Measurement and Mitigation
by: Chien, Jennifer, et al.
Published: (2024)
by: Chien, Jennifer, et al.
Published: (2024)
Generative AI and Power Imbalances in Global Education: Frameworks for Bias Mitigation
by: Nyaaba, Matthew, et al.
Published: (2024)
by: Nyaaba, Matthew, et al.
Published: (2024)
Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services
by: Dimino, Fabrizio, et al.
Published: (2026)
by: Dimino, Fabrizio, et al.
Published: (2026)
Survival at Any Cost? LLMs and the Choice Between Self-Preservation and Human Harm
by: Mohamadi, Alireza, et al.
Published: (2025)
by: Mohamadi, Alireza, et al.
Published: (2025)
Failing on Bias Mitigation: A Case Study on the Challenges of Fairness in Government Data
by: Bo, Hongbo, et al.
Published: (2026)
by: Bo, Hongbo, et al.
Published: (2026)
Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
by: Sharma, Vibhhu, et al.
Published: (2026)
by: Sharma, Vibhhu, et al.
Published: (2026)
Mitigating Bias for Question Answering Models by Tracking Bias Influence
by: Ma, Mingyu Derek, et al.
Published: (2023)
by: Ma, Mingyu Derek, et al.
Published: (2023)
Backdoor for Debias: Mitigating Model Bias with Backdoor Attack-based Artificial Bias
by: Wu, Shangxi, et al.
Published: (2023)
by: Wu, Shangxi, et al.
Published: (2023)
Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs
by: Cheng, Myra, et al.
Published: (2026)
by: Cheng, Myra, et al.
Published: (2026)
No Free Lunch in Language Model Bias Mitigation? Targeted Bias Reduction Can Exacerbate Unmitigated LLM Biases
by: Chand, Shireen, et al.
Published: (2025)
by: Chand, Shireen, et al.
Published: (2025)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
by: Li, Jing-Jing, et al.
Published: (2026)
by: Li, Jing-Jing, et al.
Published: (2026)
Evaluating Language Models for Harmful Manipulation
by: Akbulut, Canfer, et al.
Published: (2026)
by: Akbulut, Canfer, et al.
Published: (2026)
Measuring Machine Learning Harms from Stereotypes Requires Understanding Who Is Harmed by Which Errors in What Ways
by: Wang, Angelina, et al.
Published: (2024)
by: Wang, Angelina, et al.
Published: (2024)
Different Bias Under Different Criteria: Assessing Bias in LLMs with a Fact-Based Approach
by: Ko, Changgeon, et al.
Published: (2024)
by: Ko, Changgeon, et al.
Published: (2024)
Perceived Political Bias in LLMs Reduces Persuasive Abilities
by: DiGiuseppe, Matthew, et al.
Published: (2026)
by: DiGiuseppe, Matthew, et al.
Published: (2026)
EtiCor++: Towards Understanding Etiquettical Bias in LLMs
by: Dwivedi, Ashutosh, et al.
Published: (2025)
by: Dwivedi, Ashutosh, et al.
Published: (2025)
Revisiting Technical Bias Mitigation Strategies
by: Mahamadou, Abdoul Jalil Djiberou, et al.
Published: (2024)
by: Mahamadou, Abdoul Jalil Djiberou, et al.
Published: (2024)
Simulating Generative Social Agents via Theory-Informed Workflow Design
by: Yan, Yuwei, et al.
Published: (2025)
by: Yan, Yuwei, et al.
Published: (2025)
InsideOut: Measuring and Mitigating Insider-Outsider Bias in Interview Script Generation
by: Wan, Yixin, et al.
Published: (2025)
by: Wan, Yixin, et al.
Published: (2025)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
by: Chan, Yik Siu, et al.
Published: (2025)
by: Chan, Yik Siu, et al.
Published: (2025)
Widespread Gender and Pronoun Bias in Moral Judgments Across LLMs
by: Fernandes, Gustavo Lúcius, et al.
Published: (2026)
by: Fernandes, Gustavo Lúcius, et al.
Published: (2026)
An Agentic Evaluation Architecture for Historical Bias Detection in Educational Textbooks
by: Stefan, Gabriel, et al.
Published: (2026)
by: Stefan, Gabriel, et al.
Published: (2026)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
Sparks of Rationality: Do Reasoning LLMs Align with Human Judgment and Choice?
by: Tak, Ala N., et al.
Published: (2026)
by: Tak, Ala N., et al.
Published: (2026)
Red Lines and Grey Zones in the Fog of War: Benchmarking Legal Risk, Moral Harm, and Regional Bias in Large Language Model Military Decision-Making
by: Drinkall, Toby
Published: (2025)
by: Drinkall, Toby
Published: (2025)
Are LLMs (Really) Ideological? An IRT-based Analysis and Alignment Tool for Perceived Socio-Economic Bias in LLMs
by: Wachter, Jasmin, et al.
Published: (2025)
by: Wachter, Jasmin, et al.
Published: (2025)
Similar Items
-
Propaganda and Information Dissemination in the Russo-Ukrainian War: Natural Language Processing of Russian and Western Twitter Narratives
by: Gouliev, Zaur
Published: (2025) -
Uncovering Bias in Foundation Models: Impact, Testing, Harm, and Mitigation
by: Sun, Shuzhou, et al.
Published: (2025) -
Mitigating Gender Bias via Fostering Exploratory Thinking in LLMs
by: Wei, Kangda, et al.
Published: (2025) -
Evolution of AI in Education: Agentic Workflows
by: Kamalov, Firuz, et al.
Published: (2025) -
Echoes of AI Harms: A Human-LLM Synergistic Framework for Bias-Driven Harm Anticipation
by: Tantalaki, Nicoleta, et al.
Published: (2025)