Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Zeping, Ananiadou, Sophia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Locate-then-Merge: Neuron-Level Parameter Fusion for Mitigating Catastrophic Forgetting in Multimodal LLMs
by: Yu, Zeping, et al.
Published: (2025)
by: Yu, Zeping, et al.
Published: (2025)
Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question Answering
by: Yu, Zeping, et al.
Published: (2024)
by: Yu, Zeping, et al.
Published: (2024)
Interpreting Arithmetic Mechanism in Large Language Models through Comparative Neuron Analysis
by: Yu, Zeping, et al.
Published: (2024)
by: Yu, Zeping, et al.
Published: (2024)
Neuron-Level Knowledge Attribution in Large Language Models
by: Yu, Zeping, et al.
Published: (2023)
by: Yu, Zeping, et al.
Published: (2023)
Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language Models
by: Yu, Zeping, et al.
Published: (2025)
by: Yu, Zeping, et al.
Published: (2025)
How do Large Language Models Learn In-Context? Query and Key Matrices of In-Context Heads are Two Towers for Metric Learning
by: Yu, Zeping, et al.
Published: (2024)
by: Yu, Zeping, et al.
Published: (2024)
Disclosure and Mitigation of Gender Bias in LLMs
by: Dong, Xiangjue, et al.
Published: (2024)
by: Dong, Xiangjue, et al.
Published: (2024)
Mechanics of Bias and Reasoning: Interpreting the Impact of Chain-of-Thought Prompting on Gender Bias in LLMs
by: Pearman, Edie, et al.
Published: (2026)
by: Pearman, Edie, et al.
Published: (2026)
Emotion Detection for Misinformation: A Review
by: Liu, Zhiwei, et al.
Published: (2023)
by: Liu, Zhiwei, et al.
Published: (2023)
From n-gram to Attention: How Model Architectures Learn and Propagate Bias in Language Modeling
by: Kabir, Mohsinul, et al.
Published: (2025)
by: Kabir, Mohsinul, et al.
Published: (2025)
Mitigating Gender Bias via Fostering Exploratory Thinking in LLMs
by: Wei, Kangda, et al.
Published: (2025)
by: Wei, Kangda, et al.
Published: (2025)
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
by: Kabir, Mohsinul, et al.
Published: (2025)
by: Kabir, Mohsinul, et al.
Published: (2025)
Mitigating Gender Bias in Code Large Language Models via Model Editing
by: Qin, Zhanyue, et al.
Published: (2024)
by: Qin, Zhanyue, et al.
Published: (2024)
The power of Prompts: Evaluating and Mitigating Gender Bias in MT with LLMs
by: Sant, Aleix, et al.
Published: (2024)
by: Sant, Aleix, et al.
Published: (2024)
Religious Bias Landscape in Language and Text-to-Image Models: Analysis, Detection, and Debiasing Strategies
by: Abrar, Ajwad, et al.
Published: (2025)
by: Abrar, Ajwad, et al.
Published: (2025)
Disentangled VAD Representations via a Variational Framework for Political Stance Detection
by: Xu, Beiyu, et al.
Published: (2025)
by: Xu, Beiyu, et al.
Published: (2025)
Towards Interpretable Mental Health Analysis with Large Language Models
by: Yang, Kailai, et al.
Published: (2023)
by: Yang, Kailai, et al.
Published: (2023)
Factual consistency evaluation of summarization in the Era of large language models
by: Luo, Zheheng, et al.
Published: (2024)
by: Luo, Zheheng, et al.
Published: (2024)
The Lay Person's Guide to Biomedicine: Orchestrating Large Language Models
by: Luo, Zheheng, et al.
Published: (2024)
by: Luo, Zheheng, et al.
Published: (2024)
EmoLLMs: A Series of Emotional Large Language Models and Annotation Tools for Comprehensive Affective Analysis
by: Liu, Zhiwei, et al.
Published: (2024)
by: Liu, Zhiwei, et al.
Published: (2024)
MentaLLaMA: Interpretable Mental Health Analysis on Social Media with Large Language Models
by: Yang, Kailai, et al.
Published: (2023)
by: Yang, Kailai, et al.
Published: (2023)
The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Pre-trained Language Models
by: Liu, Yan, et al.
Published: (2024)
by: Liu, Yan, et al.
Published: (2024)
Implicit Graph, Explicit Retrieval: Towards Efficient and Interpretable Long-horizon Memory for Large Language Models
by: Zhang, Xin, et al.
Published: (2026)
by: Zhang, Xin, et al.
Published: (2026)
MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs
by: Liu, Zhiwei, et al.
Published: (2025)
by: Liu, Zhiwei, et al.
Published: (2025)
Mitigating Gender Bias in Contextual Word Embeddings
by: Yarrabelly, Navya, et al.
Published: (2024)
by: Yarrabelly, Navya, et al.
Published: (2024)
Exploring Safety Alignment Evaluation of LLMs in Chinese Mental Health Dialogues via LLM-as-Judge
by: Cai, Yunna, et al.
Published: (2025)
by: Cai, Yunna, et al.
Published: (2025)
Mitigating Copy Bias in In-Context Learning through Neuron Pruning
by: Ali, Ameen, et al.
Published: (2024)
by: Ali, Ameen, et al.
Published: (2024)
Detection, Classification, and Mitigation of Gender Bias in Large Language Models
by: Cheng, Xiaoqing, et al.
Published: (2025)
by: Cheng, Xiaoqing, et al.
Published: (2025)
Mitigating Biases for Instruction-following Language Models via Bias Neurons Elimination
by: Yang, Nakyeong, et al.
Published: (2023)
by: Yang, Nakyeong, et al.
Published: (2023)
ConspEmoLLM-v2: A robust and stable model to detect sentiment-transformed conspiracy theories
by: Liu, Zhiwei, et al.
Published: (2025)
by: Liu, Zhiwei, et al.
Published: (2025)
Rumor Detection by Multi-task Suffix Learning based on Time-series Dual Sentiments
by: Liu, Zhiwei, et al.
Published: (2025)
by: Liu, Zhiwei, et al.
Published: (2025)
Are Large Language Models True Healthcare Jacks-of-All-Trades? Benchmarking Across Health Professions Beyond Physician Exams
by: Luo, Zheheng, et al.
Published: (2024)
by: Luo, Zheheng, et al.
Published: (2024)
Exploring the Linear Subspace Hypothesis in Gender Bias Mitigation
by: Vargas, Francisco, et al.
Published: (2020)
by: Vargas, Francisco, et al.
Published: (2020)
Evaluating Gender Bias of LLMs in Making Morality Judgements
by: Bajaj, Divij, et al.
Published: (2024)
by: Bajaj, Divij, et al.
Published: (2024)
RAEmoLLM: Retrieval Augmented LLMs for Cross-Domain Misinformation Detection Using In-Context Learning Based on Emotional Information
by: Liu, Zhiwei, et al.
Published: (2024)
by: Liu, Zhiwei, et al.
Published: (2024)
Locating and Mitigating Gender Bias in Large Language Models
by: Cai, Yuchen, et al.
Published: (2024)
by: Cai, Yuchen, et al.
Published: (2024)
GenderAlign: An Alignment Dataset for Mitigating Gender Bias in Large Language Models
by: Zhang, Tao, et al.
Published: (2024)
by: Zhang, Tao, et al.
Published: (2024)
Understanding and Mitigating Cross-lingual Privacy Leakage via Language-specific and Universal Privacy Neurons
by: Dong, Wenshuo, et al.
Published: (2025)
by: Dong, Wenshuo, et al.
Published: (2025)
Overview of the NLPCC 2025 Shared Task: Gender Bias Mitigation Challenge
by: Li, Yizhi, et al.
Published: (2025)
by: Li, Yizhi, et al.
Published: (2025)
Projective Methods for Mitigating Gender Bias in Pre-trained Language Models
by: Dawkins, Hillary, et al.
Published: (2024)
by: Dawkins, Hillary, et al.
Published: (2024)
Similar Items
-
Locate-then-Merge: Neuron-Level Parameter Fusion for Mitigating Catastrophic Forgetting in Multimodal LLMs
by: Yu, Zeping, et al.
Published: (2025) -
Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question Answering
by: Yu, Zeping, et al.
Published: (2024) -
Interpreting Arithmetic Mechanism in Large Language Models through Comparative Neuron Analysis
by: Yu, Zeping, et al.
Published: (2024) -
Neuron-Level Knowledge Attribution in Large Language Models
by: Yu, Zeping, et al.
Published: (2023) -
Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language Models
by: Yu, Zeping, et al.
Published: (2025)