KnowBias: Mitigating Social Bias in LLMs via Know-Bias Neuron Enhancement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pan, Jinhao, Raj, Chahat, Mukherjee, Anjishnu, Mansouri, Sina, Wei, Bowen, Yada, Shloka, Zhu, Ziwei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Breaking Bias, Building Bridges: Evaluation and Mitigation of Social Biases in LLMs via Contact Hypothesis
von: Raj, Chahat, et al.
Veröffentlicht: (2024)
von: Raj, Chahat, et al.
Veröffentlicht: (2024)
Bias Association Discovery Framework for Open-Ended LLM Generations
von: Pan, Jinhao, et al.
Veröffentlicht: (2025)
von: Pan, Jinhao, et al.
Veröffentlicht: (2025)
What's Not Said Still Hurts: A Description-Based Evaluation Framework for Measuring Social Bias in LLMs
von: Pan, Jinhao, et al.
Veröffentlicht: (2025)
von: Pan, Jinhao, et al.
Veröffentlicht: (2025)
BiasDora: Exploring Hidden Biased Associations in Vision-Language Models
von: Raj, Chahat, et al.
Veröffentlicht: (2024)
von: Raj, Chahat, et al.
Veröffentlicht: (2024)
VIGNETTE: Socially Grounded Bias Evaluation for Vision-Language Models
von: Raj, Chahat, et al.
Veröffentlicht: (2025)
von: Raj, Chahat, et al.
Veröffentlicht: (2025)
Purdah and Patriarchy: Evaluating and Mitigating South Asian Biases in Open-Ended Multilingual LLM Generations
von: Rinki, Mamnuya, et al.
Veröffentlicht: (2025)
von: Rinki, Mamnuya, et al.
Veröffentlicht: (2025)
Talent or Luck? Evaluating Attribution Bias in Large Language Models
von: Raj, Chahat, et al.
Veröffentlicht: (2025)
von: Raj, Chahat, et al.
Veröffentlicht: (2025)
Do LLMs Know Tool Irrelevance? Demystifying Structural Alignment Bias in Tool Invocations
von: Liu, Yilong, et al.
Veröffentlicht: (2026)
von: Liu, Yilong, et al.
Veröffentlicht: (2026)
Mitigating Gender Bias via Fostering Exploratory Thinking in LLMs
von: Wei, Kangda, et al.
Veröffentlicht: (2025)
von: Wei, Kangda, et al.
Veröffentlicht: (2025)
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
von: Zhao, Kangwen, et al.
Veröffentlicht: (2025)
von: Zhao, Kangwen, et al.
Veröffentlicht: (2025)
Does Reasoning Introduce Bias? A Study of Social Bias Evaluation and Mitigation in LLM Reasoning
von: Wu, Xuyang, et al.
Veröffentlicht: (2025)
von: Wu, Xuyang, et al.
Veröffentlicht: (2025)
CogBias: Measuring and Mitigating Cognitive Bias in Large Language Models
von: Huang, Fan, et al.
Veröffentlicht: (2026)
von: Huang, Fan, et al.
Veröffentlicht: (2026)
Mitigating Biases for Instruction-following Language Models via Bias Neurons Elimination
von: Yang, Nakyeong, et al.
Veröffentlicht: (2023)
von: Yang, Nakyeong, et al.
Veröffentlicht: (2023)
Activation Steering for Bias Mitigation: An Interpretable Approach to Safer LLMs
von: Dubey, Shivam
Veröffentlicht: (2025)
von: Dubey, Shivam
Veröffentlicht: (2025)
Steering Towards Fairness: Mitigating Political Bias in LLMs
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
Mitigating Bias for Question Answering Models by Tracking Bias Influence
von: Ma, Mingyu Derek, et al.
Veröffentlicht: (2023)
von: Ma, Mingyu Derek, et al.
Veröffentlicht: (2023)
UniBias: Unveiling and Mitigating LLM Bias through Internal Attention and FFN Manipulation
von: Zhou, Hanzhang, et al.
Veröffentlicht: (2024)
von: Zhou, Hanzhang, et al.
Veröffentlicht: (2024)
Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure
von: Lamparth, Max, et al.
Veröffentlicht: (2026)
von: Lamparth, Max, et al.
Veröffentlicht: (2026)
BiasBusters: Uncovering and Mitigating Tool Selection Bias in Large Language Models
von: Blankenstein, Thierry, et al.
Veröffentlicht: (2025)
von: Blankenstein, Thierry, et al.
Veröffentlicht: (2025)
Identifying and Mitigating Social Bias Knowledge in Language Models
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024)
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024)
First, Do No Harm (With LLMs): Mitigating Racial Bias via Agentic Workflows
von: Xing, Sihao, et al.
Veröffentlicht: (2026)
von: Xing, Sihao, et al.
Veröffentlicht: (2026)
User-Assistant Bias in LLMs
von: Pan, Xu, et al.
Veröffentlicht: (2025)
von: Pan, Xu, et al.
Veröffentlicht: (2025)
From Oracle to Noisy Context: Mitigating Contextual Exposure Bias in Speech-LLMs
von: Guo, Xiaoyong, et al.
Veröffentlicht: (2026)
von: Guo, Xiaoyong, et al.
Veröffentlicht: (2026)
Know What You Know: Metacognitive Entropy Calibration for Verifiable RL Reasoning
von: Zhao, Qiannian, et al.
Veröffentlicht: (2026)
von: Zhao, Qiannian, et al.
Veröffentlicht: (2026)
Reinforcement Learning from Multi-role Debates as Feedback for Bias Mitigation in LLMs
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2024)
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2024)
An Empirical Survey of Model Merging Algorithms for Social Bias Mitigation
von: Shirafuji, Daiki, et al.
Veröffentlicht: (2025)
von: Shirafuji, Daiki, et al.
Veröffentlicht: (2025)
Relative Bias: A Comparative Framework for Quantifying Bias in LLMs
von: Arbabi, Alireza, et al.
Veröffentlicht: (2025)
von: Arbabi, Alireza, et al.
Veröffentlicht: (2025)
Mitigating Cognitive Bias in RLHF by Altering Rationality
von: Horter, Tiffany, et al.
Veröffentlicht: (2026)
von: Horter, Tiffany, et al.
Veröffentlicht: (2026)
From Bias Mitigation to Bias Negotiation: Governing Identity and Sociocultural Reasoning in Generative AI
von: Dunivin, Zackary Okun, et al.
Veröffentlicht: (2026)
von: Dunivin, Zackary Okun, et al.
Veröffentlicht: (2026)
Whither Bias Goes, I Will Go: An Integrative, Systematic Review of Algorithmic Bias Mitigation
von: Hickman, Louis, et al.
Veröffentlicht: (2024)
von: Hickman, Louis, et al.
Veröffentlicht: (2024)
Detecting and Mitigating Bias in LLMs through Knowledge Graph-Augmented Training
von: Kumar, Rajeev, et al.
Veröffentlicht: (2025)
von: Kumar, Rajeev, et al.
Veröffentlicht: (2025)
Subgroups Matter for Robust Bias Mitigation
von: Alloula, Anissa, et al.
Veröffentlicht: (2025)
von: Alloula, Anissa, et al.
Veröffentlicht: (2025)
AGR: Age Group fairness Reward for Bias Mitigation in LLMs
von: Cao, Shuirong, et al.
Veröffentlicht: (2024)
von: Cao, Shuirong, et al.
Veröffentlicht: (2024)
Shifting Perspectives: Steering Vectors for Robust Bias Mitigation in LLMs
von: Siddique, Zara, et al.
Veröffentlicht: (2025)
von: Siddique, Zara, et al.
Veröffentlicht: (2025)
Certifying Counterfactual Bias in LLMs
von: Chaudhary, Isha, et al.
Veröffentlicht: (2024)
von: Chaudhary, Isha, et al.
Veröffentlicht: (2024)
Capturing Bias Diversity in LLMs
von: Gosavi, Purva Prasad, et al.
Veröffentlicht: (2024)
von: Gosavi, Purva Prasad, et al.
Veröffentlicht: (2024)
Social Bias in LLM-Generated Code: Benchmark and Mitigation
von: Rabbi, Fazle, et al.
Veröffentlicht: (2026)
von: Rabbi, Fazle, et al.
Veröffentlicht: (2026)
Confidence-Orchestrated Self-Evolution against Uncertain LLM Feedback
von: Wei, Bowen, et al.
Veröffentlicht: (2026)
von: Wei, Bowen, et al.
Veröffentlicht: (2026)
Backdoor for Debias: Mitigating Model Bias with Backdoor Attack-based Artificial Bias
von: Wu, Shangxi, et al.
Veröffentlicht: (2023)
von: Wu, Shangxi, et al.
Veröffentlicht: (2023)
From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test
von: Dai, Xunlian, et al.
Veröffentlicht: (2025)
von: Dai, Xunlian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Breaking Bias, Building Bridges: Evaluation and Mitigation of Social Biases in LLMs via Contact Hypothesis
von: Raj, Chahat, et al.
Veröffentlicht: (2024) -
Bias Association Discovery Framework for Open-Ended LLM Generations
von: Pan, Jinhao, et al.
Veröffentlicht: (2025) -
What's Not Said Still Hurts: A Description-Based Evaluation Framework for Measuring Social Bias in LLMs
von: Pan, Jinhao, et al.
Veröffentlicht: (2025) -
BiasDora: Exploring Hidden Biased Associations in Vision-Language Models
von: Raj, Chahat, et al.
Veröffentlicht: (2024) -
VIGNETTE: Socially Grounded Bias Evaluation for Vision-Language Models
von: Raj, Chahat, et al.
Veröffentlicht: (2025)