Shifting Perspectives: Steering Vectors for Robust Bias Mitigation in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Siddique, Zara, Khalid, Irtaza, Turner, Liam D., Espinosa-Anke, Luis |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dialz: A Python Toolkit for Steering Vectors
by: Siddique, Zara, et al.
Published: (2025)
by: Siddique, Zara, et al.
Published: (2025)
Large Language and Reasoning Models are Shallow Disjunctive Reasoners
by: Khalid, Irtaza, et al.
Published: (2025)
by: Khalid, Irtaza, et al.
Published: (2025)
Steer2Adapt: Dynamically Composing Steering Vectors Elicits Efficient Adaptation of LLMs
by: Han, Pengrui, et al.
Published: (2026)
by: Han, Pengrui, et al.
Published: (2026)
Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language Models
by: Siddique, Zara, et al.
Published: (2024)
by: Siddique, Zara, et al.
Published: (2024)
Hoaxpedia: A Unified Wikipedia Hoax Articles Dataset
by: Borkakoty, Hsuvas, et al.
Published: (2024)
by: Borkakoty, Hsuvas, et al.
Published: (2024)
CHEW: A Dataset of CHanging Events in Wikipedia
by: Borkakoty, Hsuvas, et al.
Published: (2024)
by: Borkakoty, Hsuvas, et al.
Published: (2024)
AGR: Age Group fairness Reward for Bias Mitigation in LLMs
by: Cao, Shuirong, et al.
Published: (2024)
by: Cao, Shuirong, et al.
Published: (2024)
Systematic Relational Reasoning With Epistemic Graph Neural Networks
by: Khalid, Irtaza, et al.
Published: (2024)
by: Khalid, Irtaza, et al.
Published: (2024)
LangFIR: Discovering Sparse Language-Specific Features from Monolingual Data for Language Steering
by: Wong, Sing Hieng, et al.
Published: (2026)
by: Wong, Sing Hieng, et al.
Published: (2026)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
by: Chalnev, Sviatoslav, et al.
Published: (2024)
by: Chalnev, Sviatoslav, et al.
Published: (2024)
Understanding and Mitigating Dataset Corruption in LLM Steering
by: Anderson, Cullen, et al.
Published: (2026)
by: Anderson, Cullen, et al.
Published: (2026)
Steering LLMs for Formal Theorem Proving
by: Kirtania, Shashank, et al.
Published: (2025)
by: Kirtania, Shashank, et al.
Published: (2025)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
by: Siu, Vincent, et al.
Published: (2025)
by: Siu, Vincent, et al.
Published: (2025)
Mitigating Overthinking in Large Reasoning Models via Manifold Steering
by: Huang, Yao, et al.
Published: (2025)
by: Huang, Yao, et al.
Published: (2025)
Focus On This, Not That! Steering LLMs with Adaptive Feature Specification
by: Lamb, Tom A., et al.
Published: (2024)
by: Lamb, Tom A., et al.
Published: (2024)
Steering Llama 2 via Contrastive Activation Addition
by: Panickssery, Nina, et al.
Published: (2023)
by: Panickssery, Nina, et al.
Published: (2023)
Extracting Unlearned Information from LLMs with Activation Steering
by: Seyitoğlu, Atakan, et al.
Published: (2024)
by: Seyitoğlu, Atakan, et al.
Published: (2024)
Steering Towards Fairness: Mitigating Political Bias in LLMs
by: Nadeem, Afrozah, et al.
Published: (2025)
by: Nadeem, Afrozah, et al.
Published: (2025)
Understanding and Mitigating Tokenization Bias in Language Models
by: Phan, Buu, et al.
Published: (2024)
by: Phan, Buu, et al.
Published: (2024)
Understanding Unreliability of Steering Vectors in Language Models: Geometric Predictors and the Limits of Linear Approximations
by: Braun, Joschka
Published: (2026)
by: Braun, Joschka
Published: (2026)
Relative Bias: A Comparative Framework for Quantifying Bias in LLMs
by: Arbabi, Alireza, et al.
Published: (2025)
by: Arbabi, Alireza, et al.
Published: (2025)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
by: Yang, Jinming, et al.
Published: (2026)
by: Yang, Jinming, et al.
Published: (2026)
Mitigating Extrinsic Gender Bias for Bangla Classification Tasks
by: Joy, Sajib Kumar Saha, et al.
Published: (2024)
by: Joy, Sajib Kumar Saha, et al.
Published: (2024)
In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering
by: Liu, Sheng, et al.
Published: (2023)
by: Liu, Sheng, et al.
Published: (2023)
Mitigating Bias for Question Answering Models by Tracking Bias Influence
by: Ma, Mingyu Derek, et al.
Published: (2023)
by: Ma, Mingyu Derek, et al.
Published: (2023)
The Impact of Inference Acceleration on Bias of LLMs
by: Kirsten, Elisabeth, et al.
Published: (2024)
by: Kirsten, Elisabeth, et al.
Published: (2024)
Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines
by: Jørgensen, Mikkel Godsk, et al.
Published: (2026)
by: Jørgensen, Mikkel Godsk, et al.
Published: (2026)
Attention Speaks Volumes: Localizing and Mitigating Bias in Language Models
by: Adiga, Rishabh, et al.
Published: (2024)
by: Adiga, Rishabh, et al.
Published: (2024)
How to Train Your Advisor: Steering Black-Box LLMs with Advisor Models
by: Asawa, Parth, et al.
Published: (2025)
by: Asawa, Parth, et al.
Published: (2025)
Learning to Steer: Input-dependent Steering for Multimodal LLMs
by: Parekh, Jayneel, et al.
Published: (2025)
by: Parekh, Jayneel, et al.
Published: (2025)
MoESD: Mixture of Experts Stable Diffusion to Mitigate Gender Bias
by: Wang, Guorun, et al.
Published: (2024)
by: Wang, Guorun, et al.
Published: (2024)
HyperSteer: Activation Steering at Scale with Hypernetworks
by: Sun, Jiuding, et al.
Published: (2025)
by: Sun, Jiuding, et al.
Published: (2025)
Latent Phase-Shift Rollback: Inference-Time Error Correction via Residual Stream Monitoring and KV-Cache Steering
by: Gupta, Manan, et al.
Published: (2026)
by: Gupta, Manan, et al.
Published: (2026)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
Mitigating Degree Bias Adaptively with Hard-to-Learn Nodes in Graph Contrastive Learning
by: Hu, Jingyu, et al.
Published: (2025)
by: Hu, Jingyu, et al.
Published: (2025)
Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks
by: Li, Miaomiao, et al.
Published: (2025)
by: Li, Miaomiao, et al.
Published: (2025)
What is in a name? Mitigating Name Bias in Text Embeddings via Anonymization
by: Manchanda, Sahil, et al.
Published: (2025)
by: Manchanda, Sahil, et al.
Published: (2025)
Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPO
by: Zheng, Jinquan, et al.
Published: (2026)
by: Zheng, Jinquan, et al.
Published: (2026)
Mitigating Biases for Instruction-following Language Models via Bias Neurons Elimination
by: Yang, Nakyeong, et al.
Published: (2023)
by: Yang, Nakyeong, et al.
Published: (2023)
Steer Like the LLM: Activation Steering that Mimics Prompting
by: Heyman, Geert, et al.
Published: (2026)
by: Heyman, Geert, et al.
Published: (2026)
Similar Items
-
Dialz: A Python Toolkit for Steering Vectors
by: Siddique, Zara, et al.
Published: (2025) -
Large Language and Reasoning Models are Shallow Disjunctive Reasoners
by: Khalid, Irtaza, et al.
Published: (2025) -
Steer2Adapt: Dynamically Composing Steering Vectors Elicits Efficient Adaptation of LLMs
by: Han, Pengrui, et al.
Published: (2026) -
Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language Models
by: Siddique, Zara, et al.
Published: (2024) -
Hoaxpedia: A Unified Wikipedia Hoax Articles Dataset
by: Borkakoty, Hsuvas, et al.
Published: (2024)