Debiasing LLMs by Masking Unfairness-Driving Attention Heads
Fuente:
arXiv
Saved in:
| Main Authors: | Han, Tingxu, Song, Wei, Ding, Ziqi, Li, Ziming, Fang, Chunrong, Li, Yuekang, Liu, Dongfang, Chen, Zhenyu, Wang, Zhenting |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Token-Budget-Aware LLM Reasoning
by: Han, Tingxu, et al.
Published: (2024)
by: Han, Tingxu, et al.
Published: (2024)
Mutual Information Guided Backdoor Mitigation for Pre-trained Encoders
by: Han, Tingxu, et al.
Published: (2024)
by: Han, Tingxu, et al.
Published: (2024)
SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?
by: Han, Tingxu, et al.
Published: (2026)
by: Han, Tingxu, et al.
Published: (2026)
TombRaider: Entering the Vault of History to Jailbreak Large Language Models
by: Ding, Junchen, et al.
Published: (2025)
by: Ding, Junchen, et al.
Published: (2025)
General Phrase Debiaser: Debiasing Masked Language Models at a Multi-Token Level
by: Shi, Bingkang, et al.
Published: (2023)
by: Shi, Bingkang, et al.
Published: (2023)
"Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas
by: Ding, Junchen, et al.
Published: (2025)
by: Ding, Junchen, et al.
Published: (2025)
FAIRT2V: Training-Free Debiasing for Text-to-Video Diffusion Models
by: Zhong, Haonan, et al.
Published: (2026)
by: Zhong, Haonan, et al.
Published: (2026)
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
by: Chen, Yilong, et al.
Published: (2024)
by: Chen, Yilong, et al.
Published: (2024)
Unsupervised Text Style Transfer via LLMs and Attention Masking with Multi-way Interactions
by: Pan, Lei, et al.
Published: (2024)
by: Pan, Lei, et al.
Published: (2024)
Help or Hurdle? Rethinking Model Context Protocol-Augmented Large Language Models
by: Song, Wei, et al.
Published: (2025)
by: Song, Wei, et al.
Published: (2025)
Pitfalls of Conversational LLMs on News Debiasing
by: Schlicht, Ipek Baris, et al.
Published: (2024)
by: Schlicht, Ipek Baris, et al.
Published: (2024)
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
by: Lin, Gang, et al.
Published: (2026)
by: Lin, Gang, et al.
Published: (2026)
TokenSeek: Memory Efficient Fine Tuning via Instance-Aware Token Ditching
by: Zeng, Runjia, et al.
Published: (2026)
by: Zeng, Runjia, et al.
Published: (2026)
Social Debiasing for Fair Multi-modal LLMs
by: Cheng, Harry, et al.
Published: (2024)
by: Cheng, Harry, et al.
Published: (2024)
Debiasing Large Language Models via Adaptive Causal Prompting with Sketch-of-Thought
by: Li, Bowen, et al.
Published: (2026)
by: Li, Bowen, et al.
Published: (2026)
In-Context Learning State Vector with Inner and Momentum Optimization
by: Li, Dongfang, et al.
Published: (2024)
by: Li, Dongfang, et al.
Published: (2024)
UGID: Unified Graph Isomorphism for Debiasing Large Language Models
by: Ding, Zikang, et al.
Published: (2026)
by: Ding, Zikang, et al.
Published: (2026)
Position Debiasing Fine-Tuning for Causal Perception in Long-Term Dialogue
by: Fan, Shixuan, et al.
Published: (2024)
by: Fan, Shixuan, et al.
Published: (2024)
Causal-Guided Active Learning for Debiasing Large Language Models
by: Du, Li, et al.
Published: (2024)
by: Du, Li, et al.
Published: (2024)
Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads
by: Yeo, Wei Jie, et al.
Published: (2025)
by: Yeo, Wei Jie, et al.
Published: (2025)
Vision Enhancing LLMs: Empowering Multimodal Knowledge Storage and Sharing in LLMs
by: Li, Yunxin, et al.
Published: (2023)
by: Li, Yunxin, et al.
Published: (2023)
Intermittent Semi-Working Mask: A New Masking Paradigm for LLMs
by: Hu, HaoYuan, et al.
Published: (2024)
by: Hu, HaoYuan, et al.
Published: (2024)
Improving Attributed Text Generation of Large Language Models via Preference Learning
by: Li, Dongfang, et al.
Published: (2024)
by: Li, Dongfang, et al.
Published: (2024)
PromptBridge: Cross-Model Prompt Transfer for Large Language Models
by: Wang, Yaxuan, et al.
Published: (2025)
by: Wang, Yaxuan, et al.
Published: (2025)
On the Role of Attention Heads in Large Language Model Safety
by: Zhou, Zhenhong, et al.
Published: (2024)
by: Zhou, Zhenhong, et al.
Published: (2024)
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
by: Wang, Hanrui, et al.
Published: (2020)
by: Wang, Hanrui, et al.
Published: (2020)
Information Gain-Guided Causal Intervention for Autonomous Debiasing Large Language Models
by: Sun, Zhouhao, et al.
Published: (2025)
by: Sun, Zhouhao, et al.
Published: (2025)
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
by: Lu, Yi, et al.
Published: (2024)
by: Lu, Yi, et al.
Published: (2024)
Out-of-Distribution Detection with Attention Head Masking for Multimodal Document Classification
by: Constantinou, Christos, et al.
Published: (2024)
by: Constantinou, Christos, et al.
Published: (2024)
A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation (GALI)
by: Li, Yan, et al.
Published: (2025)
by: Li, Yan, et al.
Published: (2025)
Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs
by: Ji, Tao, et al.
Published: (2025)
by: Ji, Tao, et al.
Published: (2025)
Debiasing Multilingual LLMs in Cross-lingual Latent Space
by: Peng, Qiwei, et al.
Published: (2025)
by: Peng, Qiwei, et al.
Published: (2025)
Does RAG Introduce Unfairness in LLMs? Evaluating Fairness in Retrieval-Augmented Generation Systems
by: Wu, Xuyang, et al.
Published: (2024)
by: Wu, Xuyang, et al.
Published: (2024)
Fair Generation without Unfair Distortions: Debiasing Text-to-Image Generation with Entanglement-Free Attention
by: Park, Jeonghoon, et al.
Published: (2025)
by: Park, Jeonghoon, et al.
Published: (2025)
Semantic Refinement with LLMs for Graph Representations
by: Thapaliya, Safal, et al.
Published: (2025)
by: Thapaliya, Safal, et al.
Published: (2025)
Potential and Challenges of Model Editing for Social Debiasing
by: Yan, Jianhao, et al.
Published: (2024)
by: Yan, Jianhao, et al.
Published: (2024)
CatRAG: Functor-Guided Structural Debiasing with Retrieval Augmentation for Fair LLMs
by: Ranjan, Ravi, et al.
Published: (2026)
by: Ranjan, Ravi, et al.
Published: (2026)
Continuous Concepts Removal in Text-to-image Diffusion Models
by: Han, Tingxu, et al.
Published: (2024)
by: Han, Tingxu, et al.
Published: (2024)
Focusing on Language: Revealing and Exploiting Language Attention Heads in Multilingual Large Language Models
by: Liu, Xin, et al.
Published: (2025)
by: Liu, Xin, et al.
Published: (2025)
GaLore$+$: Boosting Low-Rank Adaptation for LLMs with Cross-Head Projection
by: Liao, Xutao, et al.
Published: (2024)
by: Liao, Xutao, et al.
Published: (2024)
Similar Items
-
Token-Budget-Aware LLM Reasoning
by: Han, Tingxu, et al.
Published: (2024) -
Mutual Information Guided Backdoor Mitigation for Pre-trained Encoders
by: Han, Tingxu, et al.
Published: (2024) -
SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?
by: Han, Tingxu, et al.
Published: (2026) -
TombRaider: Entering the Vault of History to Jailbreak Large Language Models
by: Ding, Junchen, et al.
Published: (2025) -
General Phrase Debiaser: Debiasing Masked Language Models at a Multi-Token Level
by: Shi, Bingkang, et al.
Published: (2023)