Unsupervised Concept Vector Extraction for Bias Control in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Cyberey, Hannah, Ji, Yangfeng, Evans, David |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do Prevalent Bias Metrics Capture Allocational Harms from LLMs?
by: Cyberey, Hannah, et al.
Published: (2024)
by: Cyberey, Hannah, et al.
Published: (2024)
White-Box Sensitivity Auditing with Steering Vectors
by: Cyberey, Hannah, et al.
Published: (2026)
by: Cyberey, Hannah, et al.
Published: (2026)
Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control
by: Cyberey, Hannah, et al.
Published: (2025)
by: Cyberey, Hannah, et al.
Published: (2025)
Addressing Both Statistical and Causal Gender Fairness in NLP Models
by: Chen, Hannah, et al.
Published: (2024)
by: Chen, Hannah, et al.
Published: (2024)
LLMs are Biased Teachers: Evaluating LLM Bias in Personalized Education
by: Weissburg, Iain, et al.
Published: (2024)
by: Weissburg, Iain, et al.
Published: (2024)
Dutch Metaphor Extraction from Cancer Patients' Interviews and Forum Data using LLMs and Human in the Loop
by: Han, Lifeng, et al.
Published: (2025)
by: Han, Lifeng, et al.
Published: (2025)
Are Social Sentiments Inherent in LLMs? An Empirical Study on Extraction of Inter-demographic Sentiments
by: Tanaka, Kunitomo, et al.
Published: (2024)
by: Tanaka, Kunitomo, et al.
Published: (2024)
Hidden Persuaders: LLMs' Political Leaning and Their Influence on Voters
by: Potter, Yujin, et al.
Published: (2024)
by: Potter, Yujin, et al.
Published: (2024)
Different Bias Under Different Criteria: Assessing Bias in LLMs with a Fact-Based Approach
by: Ko, Changgeon, et al.
Published: (2024)
by: Ko, Changgeon, et al.
Published: (2024)
EtiCor++: Towards Understanding Etiquettical Bias in LLMs
by: Dwivedi, Ashutosh, et al.
Published: (2025)
by: Dwivedi, Ashutosh, et al.
Published: (2025)
Perceived Political Bias in LLMs Reduces Persuasive Abilities
by: DiGiuseppe, Matthew, et al.
Published: (2026)
by: DiGiuseppe, Matthew, et al.
Published: (2026)
Sometimes the Model doth Preach: Quantifying Religious Bias in Open LLMs through Demographic Analysis in Asian Nations
by: Shankar, Hari, et al.
Published: (2025)
by: Shankar, Hari, et al.
Published: (2025)
Mitigating Gender Bias via Fostering Exploratory Thinking in LLMs
by: Wei, Kangda, et al.
Published: (2025)
by: Wei, Kangda, et al.
Published: (2025)
Widespread Gender and Pronoun Bias in Moral Judgments Across LLMs
by: Fernandes, Gustavo Lúcius, et al.
Published: (2026)
by: Fernandes, Gustavo Lúcius, et al.
Published: (2026)
Are LLMs (Really) Ideological? An IRT-based Analysis and Alignment Tool for Perceived Socio-Economic Bias in LLMs
by: Wachter, Jasmin, et al.
Published: (2025)
by: Wachter, Jasmin, et al.
Published: (2025)
A Few Good Clauses: Comparing LLMs vs Domain-Trained Small Language Models on Structured Contract Extraction
by: Lincoln, Nicole, et al.
Published: (2026)
by: Lincoln, Nicole, et al.
Published: (2026)
Counterfactual Probing for the Influence of Affect and Specificity on Intergroup Bias
by: Govindarajan, Venkata S, et al.
Published: (2023)
by: Govindarajan, Venkata S, et al.
Published: (2023)
Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs
by: Islam, Tunazzina
Published: (2026)
by: Islam, Tunazzina
Published: (2026)
Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)
by: Neumann, Anna, et al.
Published: (2025)
by: Neumann, Anna, et al.
Published: (2025)
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
by: Majumdar, Ayan, et al.
Published: (2025)
by: Majumdar, Ayan, et al.
Published: (2025)
AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic Assessment
by: Li, Kun, et al.
Published: (2025)
by: Li, Kun, et al.
Published: (2025)
Gender Bias in LLMs: Preliminary Evidence from Shared Parenting Scenario in Czech Family Law
by: Harasta, Jakub, et al.
Published: (2026)
by: Harasta, Jakub, et al.
Published: (2026)
Identifying Emerging Concepts in Large Corpora
by: Ma, Sibo, et al.
Published: (2025)
by: Ma, Sibo, et al.
Published: (2025)
Only a Little to the Left: A Theory-grounded Measure of Political Bias in Large Language Models
by: Faulborn, Mats, et al.
Published: (2025)
by: Faulborn, Mats, et al.
Published: (2025)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
by: Rystrøm, Jonathan, et al.
Published: (2025)
by: Rystrøm, Jonathan, et al.
Published: (2025)
On The Conceptualization and Societal Impact of Cross-Cultural Bias
by: Bhandari, Vitthal
Published: (2025)
by: Bhandari, Vitthal
Published: (2025)
Ask LLMs Directly, "What shapes your bias?": Measuring Social Bias in Large Language Models
by: Shin, Jisu, et al.
Published: (2024)
by: Shin, Jisu, et al.
Published: (2024)
A Comprehensive Survey of Bias in LLMs: Current Landscape and Future Directions
by: Ranjan, Rajesh, et al.
Published: (2024)
by: Ranjan, Rajesh, et al.
Published: (2024)
Quantitative Information Extraction from Humanitarian Documents
by: Liberatore, Daniele, et al.
Published: (2024)
by: Liberatore, Daniele, et al.
Published: (2024)
The Statistical Signature of LLMs
by: Hadad, Ortal, et al.
Published: (2026)
by: Hadad, Ortal, et al.
Published: (2026)
Characterizing Selective Refusal Bias in Large Language Models
by: Khorramrouz, Adel, et al.
Published: (2025)
by: Khorramrouz, Adel, et al.
Published: (2025)
Gender Bias in Emotion Recognition by Large Language Models
by: Herbert, Maureen, et al.
Published: (2025)
by: Herbert, Maureen, et al.
Published: (2025)
Theories of "Sexuality" in Natural Language Processing Bias Research
by: Hobbs, Jacob
Published: (2025)
by: Hobbs, Jacob
Published: (2025)
Mitigating Social Desirability Bias in Random Silicon Sampling
by: Chapala, Sashank, et al.
Published: (2025)
by: Chapala, Sashank, et al.
Published: (2025)
In-Context Learning (and Unlearning) of Length Biases
by: Schoch, Stephanie, et al.
Published: (2025)
by: Schoch, Stephanie, et al.
Published: (2025)
Monte Carlo Sampling for Analyzing In-Context Examples
by: Schoch, Stephanie, et al.
Published: (2025)
by: Schoch, Stephanie, et al.
Published: (2025)
A Comparative Study of Learning Paradigms in Large Language Models via Intrinsic Dimension
by: Janapati, Saahith, et al.
Published: (2024)
by: Janapati, Saahith, et al.
Published: (2024)
The Political Preferences of LLMs
by: Rozado, David
Published: (2024)
by: Rozado, David
Published: (2024)
Assessing Judging Bias in Large Reasoning Models: An Empirical Study
by: Wang, Qian, et al.
Published: (2025)
by: Wang, Qian, et al.
Published: (2025)
Beyond English: Unveiling Multilingual Bias in LLM Copyright Compliance
by: Chen, Yupeng, et al.
Published: (2025)
by: Chen, Yupeng, et al.
Published: (2025)
Similar Items
-
Do Prevalent Bias Metrics Capture Allocational Harms from LLMs?
by: Cyberey, Hannah, et al.
Published: (2024) -
White-Box Sensitivity Auditing with Steering Vectors
by: Cyberey, Hannah, et al.
Published: (2026) -
Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control
by: Cyberey, Hannah, et al.
Published: (2025) -
Addressing Both Statistical and Causal Gender Fairness in NLP Models
by: Chen, Hannah, et al.
Published: (2024) -
LLMs are Biased Teachers: Evaluating LLM Bias in Personalized Education
by: Weissburg, Iain, et al.
Published: (2024)