UniBias: Unveiling and Mitigating LLM Bias through Internal Attention and FFN Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Hanzhang, Feng, Zijian, Zhu, Zixiao, Qian, Junlang, Mao, Kezhi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unveiling and Manipulating Prompt Influence in Large Language Models
by: Feng, Zijian, et al.
Published: (2024)
by: Feng, Zijian, et al.
Published: (2024)
LLMs Learn Task Heuristics from Demonstrations: A Heuristic-Driven Prompting Strategy for Document-Level Event Argument Extraction
by: Zhou, Hanzhang, et al.
Published: (2023)
by: Zhou, Hanzhang, et al.
Published: (2023)
Logit Separability-Driven Samples and Multiple Class-Related Words Selection for Advancing In-Context Learning
by: Zixiao, Zhu, et al.
Published: (2024)
by: Zixiao, Zhu, et al.
Published: (2024)
Beyond the Next Token: Towards Prompt-Robust Zero-Shot Classification via Efficient Multi-Token Prediction
by: Qian, Junlang, et al.
Published: (2025)
by: Qian, Junlang, et al.
Published: (2025)
FreeCtrl: Constructing Control Centers with Feedforward Layers for Learning-Free Controllable Text Generation
by: Feng, Zijian, et al.
Published: (2024)
by: Feng, Zijian, et al.
Published: (2024)
Does Reasoning Introduce Bias? A Study of Social Bias Evaluation and Mitigation in LLM Reasoning
by: Wu, Xuyang, et al.
Published: (2025)
by: Wu, Xuyang, et al.
Published: (2025)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
by: Yang, Jinming, et al.
Published: (2026)
by: Yang, Jinming, et al.
Published: (2026)
Mitigating Length Bias in RLHF through a Causal Lens
by: Kim, Hyeonji, et al.
Published: (2025)
by: Kim, Hyeonji, et al.
Published: (2025)
LLM-Guided Synthetic Augmentation (LGSA) for Mitigating Bias in AI Systems
by: Karri, Sai Suhruth Reddy, et al.
Published: (2025)
by: Karri, Sai Suhruth Reddy, et al.
Published: (2025)
Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge
by: Fujinuma, Yoshinari
Published: (2025)
by: Fujinuma, Yoshinari
Published: (2025)
Attention Speaks Volumes: Localizing and Mitigating Bias in Language Models
by: Adiga, Rishabh, et al.
Published: (2024)
by: Adiga, Rishabh, et al.
Published: (2024)
Detecting and Mitigating Bias in LLMs through Knowledge Graph-Augmented Training
by: Kumar, Rajeev, et al.
Published: (2025)
by: Kumar, Rajeev, et al.
Published: (2025)
No Free Lunch in Language Model Bias Mitigation? Targeted Bias Reduction Can Exacerbate Unmitigated LLM Biases
by: Chand, Shireen, et al.
Published: (2025)
by: Chand, Shireen, et al.
Published: (2025)
Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks
by: Li, Miaomiao, et al.
Published: (2025)
by: Li, Miaomiao, et al.
Published: (2025)
Identifying and Mitigating Social Bias Knowledge in Language Models
by: Chen, Ruizhe, et al.
Published: (2024)
by: Chen, Ruizhe, et al.
Published: (2024)
Fine-Grained Activation Steering: Steering Less, Achieving More
by: Feng, Zijian, et al.
Published: (2026)
by: Feng, Zijian, et al.
Published: (2026)
Mitigation of Gender and Ethnicity Bias in AI-Generated Stories through Model Explanations
by: Dimgba, Martha O., et al.
Published: (2025)
by: Dimgba, Martha O., et al.
Published: (2025)
Mitigating Translationese Bias in Multilingual LLM-as-a-Judge via Disentangled Information Bottleneck
by: Zhang, Hongbin, et al.
Published: (2026)
by: Zhang, Hongbin, et al.
Published: (2026)
Domain Lexical Knowledge-based Word Embedding Learning for Text Classification under Small Data
by: Zhu, Zixiao, et al.
Published: (2025)
by: Zhu, Zixiao, et al.
Published: (2025)
Alleviating Choice Supportive Bias in LLM with Reasoning Dependency Generation
by: Zhuang, Nan, et al.
Published: (2025)
by: Zhuang, Nan, et al.
Published: (2025)
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
by: Demchak, Nathaniel, et al.
Published: (2024)
by: Demchak, Nathaniel, et al.
Published: (2024)
Mitigating Bias for Question Answering Models by Tracking Bias Influence
by: Ma, Mingyu Derek, et al.
Published: (2023)
by: Ma, Mingyu Derek, et al.
Published: (2023)
Unveiling the "Fairness Seesaw": Discovering and Mitigating Gender and Race Bias in Vision-Language Models
by: Lan, Jian, et al.
Published: (2025)
by: Lan, Jian, et al.
Published: (2025)
Steering Towards Fairness: Mitigating Political Bias in LLMs
by: Nadeem, Afrozah, et al.
Published: (2025)
by: Nadeem, Afrozah, et al.
Published: (2025)
Locating and Mitigating Gender Bias in Large Language Models
by: Cai, Yuchen, et al.
Published: (2024)
by: Cai, Yuchen, et al.
Published: (2024)
Mitigating Metric Bias in Minimum Bayes Risk Decoding
by: Kovacs, Geza, et al.
Published: (2024)
by: Kovacs, Geza, et al.
Published: (2024)
Correcting Negative Bias in Large Language Models through Negative Attention Score Alignment
by: Yu, Sangwon, et al.
Published: (2024)
by: Yu, Sangwon, et al.
Published: (2024)
Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement
by: Xu, Wenda, et al.
Published: (2024)
by: Xu, Wenda, et al.
Published: (2024)
No LLM is Free From Bias: A Comprehensive Study of Bias Evaluation in Large Language Models
by: Kumar, Charaka Vinayak, et al.
Published: (2025)
by: Kumar, Charaka Vinayak, et al.
Published: (2025)
Likelihood-based Mitigation of Evaluation Bias in Large Language Models
by: Oi, Masanari, et al.
Published: (2024)
by: Oi, Masanari, et al.
Published: (2024)
Multi-Persona Thinking for Bias Mitigation in Large Language Models
by: Chen, Yuxing, et al.
Published: (2026)
by: Chen, Yuxing, et al.
Published: (2026)
An Empirical Survey of Model Merging Algorithms for Social Bias Mitigation
by: Shirafuji, Daiki, et al.
Published: (2025)
by: Shirafuji, Daiki, et al.
Published: (2025)
A Variational Approach for Mitigating Entity Bias in Relation Extraction
by: Mensah, Samuel, et al.
Published: (2025)
by: Mensah, Samuel, et al.
Published: (2025)
Equilibrium Dynamics and Mitigation of Gender Bias in Synthetically Generated Data
by: Kattamuri, Ashish, et al.
Published: (2025)
by: Kattamuri, Ashish, et al.
Published: (2025)
Gender Bias in LLM-generated Interview Responses
by: Kong, Haein, et al.
Published: (2024)
by: Kong, Haein, et al.
Published: (2024)
From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test
by: Dai, Xunlian, et al.
Published: (2025)
by: Dai, Xunlian, et al.
Published: (2025)
Enhancing Diagnostic Accuracy through Multi-Agent Conversations: Using Large Language Models to Mitigate Cognitive Bias
by: Ke, Yu He, et al.
Published: (2024)
by: Ke, Yu He, et al.
Published: (2024)
Understanding and Mitigating Tokenization Bias in Language Models
by: Phan, Buu, et al.
Published: (2024)
by: Phan, Buu, et al.
Published: (2024)
Fairness or Fluency? An Investigation into Language Bias of Pairwise LLM-as-a-Judge
by: Zhou, Xiaolin, et al.
Published: (2026)
by: Zhou, Xiaolin, et al.
Published: (2026)
Large Language Model Bias Mitigation from the Perspective of Knowledge Editing
by: Chen, Ruizhe, et al.
Published: (2024)
by: Chen, Ruizhe, et al.
Published: (2024)
Similar Items
-
Unveiling and Manipulating Prompt Influence in Large Language Models
by: Feng, Zijian, et al.
Published: (2024) -
LLMs Learn Task Heuristics from Demonstrations: A Heuristic-Driven Prompting Strategy for Document-Level Event Argument Extraction
by: Zhou, Hanzhang, et al.
Published: (2023) -
Logit Separability-Driven Samples and Multiple Class-Related Words Selection for Advancing In-Context Learning
by: Zixiao, Zhu, et al.
Published: (2024) -
Beyond the Next Token: Towards Prompt-Robust Zero-Shot Classification via Efficient Multi-Token Prediction
by: Qian, Junlang, et al.
Published: (2025) -
FreeCtrl: Constructing Control Centers with Feedforward Layers for Learning-Free Controllable Text Generation
by: Feng, Zijian, et al.
Published: (2024)