Mitigating Biases for Instruction-following Language Models via Bias Neurons Elimination
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Nakyeong, Kang, Taegwan, Choi, Jungkyu, Lee, Honglak, Jung, Kyomin |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Wording Steers the Evaluation: Framing Bias in LLM judges
by: Hwang, Yerin, et al.
Published: (2026)
by: Hwang, Yerin, et al.
Published: (2026)
LongStory: Coherent, Complete and Length Controlled Long story Generation
by: Park, Kyeongman, et al.
Published: (2023)
by: Park, Kyeongman, et al.
Published: (2023)
Avoidance Decoding for Diverse Multi-Branch Story Generation
by: Park, Kyeongman, et al.
Published: (2025)
by: Park, Kyeongman, et al.
Published: (2025)
LLMs can be easily Confused by Instructional Distractions
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
Rethinking Post-Unlearning Behavior of Large Vision-Language Models
by: Kim, Minsung, et al.
Published: (2025)
by: Kim, Minsung, et al.
Published: (2025)
Unplug and Play Language Models: Decomposing Experts in Language Models at Inference Time
by: Yang, Nakyeong, et al.
Published: (2024)
by: Yang, Nakyeong, et al.
Published: (2024)
FaithUn: Toward Faithful Forgetting in Language Models by Investigating the Interconnectedness of Knowledge
by: Yang, Nakyeong, et al.
Published: (2025)
by: Yang, Nakyeong, et al.
Published: (2025)
How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models
by: Kim, Minsung, et al.
Published: (2025)
by: Kim, Minsung, et al.
Published: (2025)
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
by: Min, Kyungmin, et al.
Published: (2024)
by: Min, Kyungmin, et al.
Published: (2024)
KL Penalty Control via Perturbation for Direct Preference Optimization
by: Lee, Sangkyu, et al.
Published: (2025)
by: Lee, Sangkyu, et al.
Published: (2025)
MVMR: A New Framework for Evaluating Faithfulness of Video Moment Retrieval against Multiple Distractors
by: Yang, Nakyeong, et al.
Published: (2023)
by: Yang, Nakyeong, et al.
Published: (2023)
BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models
by: Lee, Isack, et al.
Published: (2024)
by: Lee, Isack, et al.
Published: (2024)
Persona is a Double-edged Sword: Mitigating the Negative Impact of Role-playing Prompts in Zero-shot Reasoning Tasks
by: Kim, Junseok, et al.
Published: (2024)
by: Kim, Junseok, et al.
Published: (2024)
Reliability-Aware Adaptive Self-Consistency for Efficient Sampling in LLM Reasoning
by: Kim, Junseok, et al.
Published: (2026)
by: Kim, Junseok, et al.
Published: (2026)
Bias Vector: Mitigating Biases in Language Models with Task Arithmetic Approach
by: Shirafuji, Daiki, et al.
Published: (2024)
by: Shirafuji, Daiki, et al.
Published: (2024)
Understanding and Mitigating Tokenization Bias in Language Models
by: Phan, Buu, et al.
Published: (2024)
by: Phan, Buu, et al.
Published: (2024)
Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPO
by: Zheng, Jinquan, et al.
Published: (2026)
by: Zheng, Jinquan, et al.
Published: (2026)
Auto-Intent: Automated Intent Discovery and Self-Exploration for Large Language Model Web Agents
by: Kim, Jaekyeom, et al.
Published: (2024)
by: Kim, Jaekyeom, et al.
Published: (2024)
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
by: Moon, Jiwon, et al.
Published: (2025)
by: Moon, Jiwon, et al.
Published: (2025)
Conditional [MASK] Discrete Diffusion Language Model
by: Koh, Hyukhun, et al.
Published: (2024)
by: Koh, Hyukhun, et al.
Published: (2024)
Attention Speaks Volumes: Localizing and Mitigating Bias in Language Models
by: Adiga, Rishabh, et al.
Published: (2024)
by: Adiga, Rishabh, et al.
Published: (2024)
FairPy: A Toolkit for Evaluation of Prediction Biases and their Mitigation in Large Language Models
by: Viswanath, Hrishikesh, et al.
Published: (2023)
by: Viswanath, Hrishikesh, et al.
Published: (2023)
ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection
by: Kim, Jeonghye, et al.
Published: (2025)
by: Kim, Jeonghye, et al.
Published: (2025)
EXAONE 3.0 7.8B Instruction Tuned Language Model
by: An, Soyoung, et al.
Published: (2024)
by: An, Soyoung, et al.
Published: (2024)
CPR: Mitigating Large Language Model Hallucinations with Curative Prompt Refinement
by: Shim, Jung-Woo, et al.
Published: (2025)
by: Shim, Jung-Woo, et al.
Published: (2025)
Multi-stage Prompt Refinement for Mitigating Hallucinations in Large Language Models
by: Shim, Jung-Woo, et al.
Published: (2025)
by: Shim, Jung-Woo, et al.
Published: (2025)
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
Instruction Mining: Instruction Data Selection for Tuning Large Language Models
by: Cao, Yihan, et al.
Published: (2023)
by: Cao, Yihan, et al.
Published: (2023)
Persona Switch: Mixing Distinct Perspectives in Decoding Time
by: Kim, Junseok, et al.
Published: (2026)
by: Kim, Junseok, et al.
Published: (2026)
Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models
by: Dong, Guanting, et al.
Published: (2024)
by: Dong, Guanting, et al.
Published: (2024)
Benchmarking Cognitive Biases in Large Language Models as Evaluators
by: Koo, Ryan, et al.
Published: (2023)
by: Koo, Ryan, et al.
Published: (2023)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
No Free Lunch in Language Model Bias Mitigation? Targeted Bias Reduction Can Exacerbate Unmitigated LLM Biases
by: Chand, Shireen, et al.
Published: (2025)
by: Chand, Shireen, et al.
Published: (2025)
CROW: Eliminating Backdoors from Large Language Models via Internal Consistency Regularization
by: Min, Nay Myat, et al.
Published: (2024)
by: Min, Nay Myat, et al.
Published: (2024)
Program Synthesis via Test-Time Transduction
by: Lee, Kang-il, et al.
Published: (2025)
by: Lee, Kang-il, et al.
Published: (2025)
Reasoning Towards Fairness: Mitigating Bias in Language Models through Reasoning-Guided Fine-Tuning
by: Kabra, Sanchit, et al.
Published: (2025)
by: Kabra, Sanchit, et al.
Published: (2025)
Confidence Regulation Neurons in Language Models
by: Stolfo, Alessandro, et al.
Published: (2024)
by: Stolfo, Alessandro, et al.
Published: (2024)
Relative Value Biases in Large Language Models
by: Hayes, William M., et al.
Published: (2024)
by: Hayes, William M., et al.
Published: (2024)
Large Language Models are Biased Reinforcement Learners
by: Hayes, William M., et al.
Published: (2024)
by: Hayes, William M., et al.
Published: (2024)
Identifying Features Associated with Bias Against 93 Stigmatized Groups in Language Models and Guardrail Model Safety Mitigation
by: Gueorguieva, Anna-Maria, et al.
Published: (2025)
by: Gueorguieva, Anna-Maria, et al.
Published: (2025)
Similar Items
-
When Wording Steers the Evaluation: Framing Bias in LLM judges
by: Hwang, Yerin, et al.
Published: (2026) -
LongStory: Coherent, Complete and Length Controlled Long story Generation
by: Park, Kyeongman, et al.
Published: (2023) -
Avoidance Decoding for Diverse Multi-Branch Story Generation
by: Park, Kyeongman, et al.
Published: (2025) -
LLMs can be easily Confused by Instructional Distractions
by: Hwang, Yerin, et al.
Published: (2025) -
Rethinking Post-Unlearning Behavior of Large Vision-Language Models
by: Kim, Minsung, et al.
Published: (2025)