White-Box Sensitivity Auditing with Steering Vectors
Fuente:
arXiv
Saved in:
| Main Authors: | Cyberey, Hannah, Ji, Yangfeng, Evans, David |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unsupervised Concept Vector Extraction for Bias Control in LLMs
by: Cyberey, Hannah, et al.
Published: (2025)
by: Cyberey, Hannah, et al.
Published: (2025)
Do Prevalent Bias Metrics Capture Allocational Harms from LLMs?
by: Cyberey, Hannah, et al.
Published: (2024)
by: Cyberey, Hannah, et al.
Published: (2024)
Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control
by: Cyberey, Hannah, et al.
Published: (2025)
by: Cyberey, Hannah, et al.
Published: (2025)
Addressing Both Statistical and Causal Gender Fairness in NLP Models
by: Chen, Hannah, et al.
Published: (2024)
by: Chen, Hannah, et al.
Published: (2024)
Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs
by: Hartmann, David, et al.
Published: (2026)
by: Hartmann, David, et al.
Published: (2026)
DSO: Direct Steering Optimization for Bias Mitigation
by: Paes, Lucas Monteiro, et al.
Published: (2025)
by: Paes, Lucas Monteiro, et al.
Published: (2025)
Representation Surgery: Theory and Practice of Affine Steering
by: Singh, Shashwat, et al.
Published: (2024)
by: Singh, Shashwat, et al.
Published: (2024)
Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive
by: Weerasooriya, Tharindu Cyril, et al.
Published: (2023)
by: Weerasooriya, Tharindu Cyril, et al.
Published: (2023)
Watching the Watchers: A Comparative Fairness Audit of Cloud-based Content Moderation Services
by: Hartmann, David, et al.
Published: (2024)
by: Hartmann, David, et al.
Published: (2024)
Group Fairness Meets the Black Box: Enabling Fair Algorithms on Closed LLMs via Post-Processing
by: Xian, Ruicheng, et al.
Published: (2025)
by: Xian, Ruicheng, et al.
Published: (2025)
Activation Steering via Generative Causal Mediation
by: Sankaranarayanan, Aruna, et al.
Published: (2026)
by: Sankaranarayanan, Aruna, et al.
Published: (2026)
Extending Activation Steering to Broad Skills and Multiple Behaviours
by: van der Weij, Teun, et al.
Published: (2024)
by: van der Weij, Teun, et al.
Published: (2024)
Reasoning Models Generate Societies of Thought
by: Kim, Junsol, et al.
Published: (2026)
by: Kim, Junsol, et al.
Published: (2026)
What's in a Name? Auditing Large Language Models for Race and Gender Bias
by: Salinas, Alejandro, et al.
Published: (2024)
by: Salinas, Alejandro, et al.
Published: (2024)
Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text
by: Islam, Tunazzina
Published: (2026)
by: Islam, Tunazzina
Published: (2026)
Predicting Where Steering Vectors Succeed
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
KPoEM: A Human-Annotated Dataset for Emotion Classification and RAG-Based Poetry Generation in Korean Modern Poetry
by: Lim, Iro, et al.
Published: (2025)
by: Lim, Iro, et al.
Published: (2025)
The "Colonial Impulse" of Natural Language Processing: An Audit of Bengali Sentiment Analysis Tools and Their Identity-based Biases
by: Das, Dipto, et al.
Published: (2024)
by: Das, Dipto, et al.
Published: (2024)
Beyond Steering Vector: Flow-based Activation Steering for Inference-Time Intervention
by: Jin, Zehao, et al.
Published: (2026)
by: Jin, Zehao, et al.
Published: (2026)
Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the U.S. Workforce
by: Shao, Yijia, et al.
Published: (2025)
by: Shao, Yijia, et al.
Published: (2025)
Thinking Outside the (Gray) Box: A Context-Based Score for Assessing Value and Originality in Neural Text Generation
by: Franceschelli, Giorgio, et al.
Published: (2025)
by: Franceschelli, Giorgio, et al.
Published: (2025)
Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
by: Girrbach, Leander, et al.
Published: (2025)
by: Girrbach, Leander, et al.
Published: (2025)
Leveraging Imperfect Sources to Detect Fairwashing in Black-Box Auditing
by: Bourrée, Jade Garcia, et al.
Published: (2023)
by: Bourrée, Jade Garcia, et al.
Published: (2023)
Semantic Sensitivities and Inconsistent Predictions: Measuring the Fragility of NLI Models
by: Arakelyan, Erik, et al.
Published: (2024)
by: Arakelyan, Erik, et al.
Published: (2024)
VSPO: Vector-Steered Policy Optimization for Behavioral Control
by: Zhang, Xuechen, et al.
Published: (2026)
by: Zhang, Xuechen, et al.
Published: (2026)
Beyond Multiple Choice: Evaluating Steering Vectors for Summarization
by: Braun, Joschka, et al.
Published: (2025)
by: Braun, Joschka, et al.
Published: (2025)
Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization
by: Cao, Yuanpu, et al.
Published: (2024)
by: Cao, Yuanpu, et al.
Published: (2024)
The Moral Foundations Reddit Corpus
by: Trager, Jackson, et al.
Published: (2022)
by: Trager, Jackson, et al.
Published: (2022)
Toward Preference-aligned Large Language Models via Residual-based Model Steering
by: La Cava, Lucio, et al.
Published: (2025)
by: La Cava, Lucio, et al.
Published: (2025)
AI-Mediated Communication Can Steer Collective Opinion
by: Tsirtsis, Stratis, et al.
Published: (2026)
by: Tsirtsis, Stratis, et al.
Published: (2026)
Protected group bias and stereotypes in Large Language Models
by: Kotek, Hadas, et al.
Published: (2024)
by: Kotek, Hadas, et al.
Published: (2024)
Toward Inclusive Educational AI: Auditing Frontier LLMs through a Multiplexity Lens
by: Mushtaq, Abdullah, et al.
Published: (2025)
by: Mushtaq, Abdullah, et al.
Published: (2025)
LEACE: Perfect linear concept erasure in closed form
by: Belrose, Nora, et al.
Published: (2023)
by: Belrose, Nora, et al.
Published: (2023)
RSD: A Local Triangulation Audit Primitive for Learned Vector Blocks
by: Jin, Seungmin
Published: (2026)
by: Jin, Seungmin
Published: (2026)
Direction-Flipped Influence Audits Reveal Hidden Structure in Moral Choices of LLMs
by: Blandfort, Phil, et al.
Published: (2026)
by: Blandfort, Phil, et al.
Published: (2026)
Linear Representations of Political Perspective Emerge in Large Language Models
by: Kim, Junsol, et al.
Published: (2025)
by: Kim, Junsol, et al.
Published: (2025)
Steer2Adapt: Dynamically Composing Steering Vectors Elicits Efficient Adaptation of LLMs
by: Han, Pengrui, et al.
Published: (2026)
by: Han, Pengrui, et al.
Published: (2026)
Auditing for Human Expertise
by: Alur, Rohan, et al.
Published: (2023)
by: Alur, Rohan, et al.
Published: (2023)
Taxonomy and Analysis of Sensitive User Queries in Generative AI Search
by: Jo, Hwiyeol, et al.
Published: (2024)
by: Jo, Hwiyeol, et al.
Published: (2024)
Reward Models Inherit Value Biases from Pretraining
by: Christian, Brian, et al.
Published: (2026)
by: Christian, Brian, et al.
Published: (2026)
Similar Items
-
Unsupervised Concept Vector Extraction for Bias Control in LLMs
by: Cyberey, Hannah, et al.
Published: (2025) -
Do Prevalent Bias Metrics Capture Allocational Harms from LLMs?
by: Cyberey, Hannah, et al.
Published: (2024) -
Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control
by: Cyberey, Hannah, et al.
Published: (2025) -
Addressing Both Statistical and Causal Gender Fairness in NLP Models
by: Chen, Hannah, et al.
Published: (2024) -
Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs
by: Hartmann, David, et al.
Published: (2026)