Towards Resource Efficient and Interpretable Bias Mitigation in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Tong, Schrasing, Zemour, Eliott, Lu, Jessica, Lohanimit, Rawisara, Kagal, Lalana |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mitigating Bias in Concept Bottleneck Models for Fair and Interpretable Image Classification
by: Tong, Schrasing, et al.
Published: (2026)
by: Tong, Schrasing, et al.
Published: (2026)
Investigating Model Editing for Unlearning in Large Language Models
by: Hossain, Shariqah, et al.
Published: (2025)
by: Hossain, Shariqah, et al.
Published: (2025)
Measuring Perceptions of Fairness in AI Systems: The Effects of Infra-marginality
by: Tong, Schrasing, et al.
Published: (2026)
by: Tong, Schrasing, et al.
Published: (2026)
Learning Concept Bottleneck Models from Mechanistic Explanations
by: De Santis, Antonio, et al.
Published: (2026)
by: De Santis, Antonio, et al.
Published: (2026)
PrimeGuard: Safe and Helpful LLMs through Tuning-Free Routing
by: Manczak, Blazej, et al.
Published: (2024)
by: Manczak, Blazej, et al.
Published: (2024)
Privacy in Image Datasets: A Case Study on Pregnancy Ultrasounds
by: Lohanimit, Rawisara, et al.
Published: (2026)
by: Lohanimit, Rawisara, et al.
Published: (2026)
Mitigating the Bias of Large Language Model Evaluation
by: Zhou, Hongli, et al.
Published: (2024)
by: Zhou, Hongli, et al.
Published: (2024)
Does fine-tuning GPT-3 with the OpenAI API leak personally-identifiable information?
by: Sun, Albert Yu, et al.
Published: (2023)
by: Sun, Albert Yu, et al.
Published: (2023)
Open-DeBias: Toward Mitigating Open-Set Bias in Language Models
by: Rani, Arti, et al.
Published: (2025)
by: Rani, Arti, et al.
Published: (2025)
Mitigating Label Length Bias in Large Language Models
by: Sanz-Guerrero, Mario, et al.
Published: (2025)
by: Sanz-Guerrero, Mario, et al.
Published: (2025)
Do Multilingual Large Language Models Mitigate Stereotype Bias?
by: Nie, Shangrui, et al.
Published: (2024)
by: Nie, Shangrui, et al.
Published: (2024)
Detection, Classification, and Mitigation of Gender Bias in Large Language Models
by: Cheng, Xiaoqing, et al.
Published: (2025)
by: Cheng, Xiaoqing, et al.
Published: (2025)
Towards Typologically Aware Rescoring to Mitigate Unfaithfulness in Lower-Resource Languages
by: Chan, Tsan Tsai, et al.
Published: (2025)
by: Chan, Tsan Tsai, et al.
Published: (2025)
Locating and Mitigating Gender Bias in Large Language Models
by: Cai, Yuchen, et al.
Published: (2024)
by: Cai, Yuchen, et al.
Published: (2024)
Bias in Large Language Models: Origin, Evaluation, and Mitigation
by: Guo, Yufei, et al.
Published: (2024)
by: Guo, Yufei, et al.
Published: (2024)
MBIAS: Mitigating Bias in Large Language Models While Retaining Context
by: Raza, Shaina, et al.
Published: (2024)
by: Raza, Shaina, et al.
Published: (2024)
Unveiling and Mitigating Bias in Mental Health Analysis with Large Language Models
by: Wang, Yuqing, et al.
Published: (2024)
by: Wang, Yuqing, et al.
Published: (2024)
Interpreting Bias in Large Language Models: A Feature-Based Approach
by: Prakash, Nirmalendu, et al.
Published: (2024)
by: Prakash, Nirmalendu, et al.
Published: (2024)
Likelihood-based Mitigation of Evaluation Bias in Large Language Models
by: Oi, Masanari, et al.
Published: (2024)
by: Oi, Masanari, et al.
Published: (2024)
Multi-Persona Thinking for Bias Mitigation in Large Language Models
by: Chen, Yuxing, et al.
Published: (2026)
by: Chen, Yuxing, et al.
Published: (2026)
Evaluating and Mitigating Social Bias for Large Language Models in Open-ended Settings
by: Liu, Zhao, et al.
Published: (2024)
by: Liu, Zhao, et al.
Published: (2024)
Mitigating Boundary Ambiguity and Inherent Bias for Text Classification in the Era of Large Language Models
by: Lu, Zhenyi, et al.
Published: (2024)
by: Lu, Zhenyi, et al.
Published: (2024)
Large Language Model Bias Mitigation from the Perspective of Knowledge Editing
by: Chen, Ruizhe, et al.
Published: (2024)
by: Chen, Ruizhe, et al.
Published: (2024)
LNE-Blocking: An Efficient Framework for Contamination Mitigation Evaluation on Large Language Models
by: Hou, Ruijie, et al.
Published: (2025)
by: Hou, Ruijie, et al.
Published: (2025)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
Faithful-Patchscopes: Understanding and Mitigating Model Bias in Hidden Representations Explanation of Large Language Models
by: Gong, Xilin, et al.
Published: (2026)
by: Gong, Xilin, et al.
Published: (2026)
Mitigating Judgment Preference Bias in Large Language Models through Group-Based Polling
by: Liu, Shuliang, et al.
Published: (2025)
by: Liu, Shuliang, et al.
Published: (2025)
FTTE: Enabling Federated and Resource-Constrained Deep Edge Intelligence
by: Tenison, Irene, et al.
Published: (2025)
by: Tenison, Irene, et al.
Published: (2025)
BiasDPO: Mitigating Bias in Language Models through Direct Preference Optimization
by: Allam, Ahmed
Published: (2024)
by: Allam, Ahmed
Published: (2024)
Simulating a Bias Mitigation Scenario in Large Language Models
by: Kiashemshaki, Kiana, et al.
Published: (2025)
by: Kiashemshaki, Kiana, et al.
Published: (2025)
Mitigating Biases in Language Models via Bias Unlearning
by: Liu, Dianqing, et al.
Published: (2025)
by: Liu, Dianqing, et al.
Published: (2025)
Language Lives in Sparse Dimensions: Toward Interpretable and Efficient Multilingual Control for Large Language Models
by: Zhong, Chengzhi, et al.
Published: (2025)
by: Zhong, Chengzhi, et al.
Published: (2025)
Towards Interpretable Mental Health Analysis with Large Language Models
by: Yang, Kailai, et al.
Published: (2023)
by: Yang, Kailai, et al.
Published: (2023)
Mitigating Social Bias in Large Language Models: A Multi-Objective Approach within a Multi-Agent Framework
by: Xu, Zhenjie, et al.
Published: (2024)
by: Xu, Zhenjie, et al.
Published: (2024)
GenderAlign: An Alignment Dataset for Mitigating Gender Bias in Large Language Models
by: Zhang, Tao, et al.
Published: (2024)
by: Zhang, Tao, et al.
Published: (2024)
Implicit Graph, Explicit Retrieval: Towards Efficient and Interpretable Long-horizon Memory for Large Language Models
by: Zhang, Xin, et al.
Published: (2026)
by: Zhang, Xin, et al.
Published: (2026)
Measuring Spiritual Values and Bias of Large Language Models
by: Liu, Songyuan, et al.
Published: (2024)
by: Liu, Songyuan, et al.
Published: (2024)
Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing
by: Yu, Zeping, et al.
Published: (2025)
by: Yu, Zeping, et al.
Published: (2025)
Context-Aware Counterfactual Data Augmentation for Gender Bias Mitigation in Language Models
by: Parihar, Shweta, et al.
Published: (2026)
by: Parihar, Shweta, et al.
Published: (2026)
Bias Attribution in Filipino Language Models: Extending a Bias Interpretability Metric for Application on Agglutinative Languages
by: Gamboa, Lance Calvin Lim, et al.
Published: (2025)
by: Gamboa, Lance Calvin Lim, et al.
Published: (2025)
Similar Items
-
Mitigating Bias in Concept Bottleneck Models for Fair and Interpretable Image Classification
by: Tong, Schrasing, et al.
Published: (2026) -
Investigating Model Editing for Unlearning in Large Language Models
by: Hossain, Shariqah, et al.
Published: (2025) -
Measuring Perceptions of Fairness in AI Systems: The Effects of Infra-marginality
by: Tong, Schrasing, et al.
Published: (2026) -
Learning Concept Bottleneck Models from Mechanistic Explanations
by: De Santis, Antonio, et al.
Published: (2026) -
PrimeGuard: Safe and Helpful LLMs through Tuning-Free Routing
by: Manczak, Blazej, et al.
Published: (2024)