Measuring Mechanistic Independence: Can Bias Be Removed Without Erasing Demographics?
Fuente:
arXiv
Saved in:
| Main Authors: | Shan, Zhengyang, Mueller, Aaron |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MABR: Multilayer Adversarial Bias Removal Without Prior Bias Knowledge
by: Yin, Maxwell J., et al.
Published: (2024)
by: Yin, Maxwell J., et al.
Published: (2024)
Obscured but Not Erased: Evaluating Nationality Bias in LLMs via Name-Based Bias Benchmarks
by: Pelosio, Giulio, et al.
Published: (2025)
by: Pelosio, Giulio, et al.
Published: (2025)
Erasing Without Remembering: Implicit Knowledge Forgetting in Large Language Models
by: Wang, Huazheng, et al.
Published: (2025)
by: Wang, Huazheng, et al.
Published: (2025)
In-Context Learning Without Copying
by: Sahin, Kerem, et al.
Published: (2025)
by: Sahin, Kerem, et al.
Published: (2025)
RedacBench: Can AI Erase Your Secrets?
by: Jeon, Hyunjun, et al.
Published: (2026)
by: Jeon, Hyunjun, et al.
Published: (2026)
Elucidating Mechanisms of Demographic Bias in LLMs for Healthcare
by: Ahsan, Hiba, et al.
Published: (2025)
by: Ahsan, Hiba, et al.
Published: (2025)
Missed Causes and Ambiguous Effects: Counterfactuals Pose Challenges for Interpreting Neural Networks
by: Mueller, Aaron
Published: (2024)
by: Mueller, Aaron
Published: (2024)
Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics
by: Nikankin, Yaniv, et al.
Published: (2024)
by: Nikankin, Yaniv, et al.
Published: (2024)
Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation
by: Cheng, Yinjie, et al.
Published: (2025)
by: Cheng, Yinjie, et al.
Published: (2025)
Assessing the Reliability of LLMs Annotations in the Context of Demographic Bias and Model Explanation
by: Mohammadi, Hadi, et al.
Published: (2025)
by: Mohammadi, Hadi, et al.
Published: (2025)
Incremental Sentence Processing Mechanisms in Autoregressive Transformer Language Models
by: Hanna, Michael, et al.
Published: (2024)
by: Hanna, Michael, et al.
Published: (2024)
Different Demographic Cues Yield Inconsistent Conclusions About LLM Personalization and Bias
by: Tonneau, Manuel, et al.
Published: (2026)
by: Tonneau, Manuel, et al.
Published: (2026)
Demographic and Linguistic Bias Evaluation in Omnimodal Language Models
by: Elobaid, Alaa
Published: (2026)
by: Elobaid, Alaa
Published: (2026)
Order-Independence Without Fine Tuning
by: McIlroy-Young, Reid, et al.
Published: (2024)
by: McIlroy-Young, Reid, et al.
Published: (2024)
Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective
by: Chandna, Bhavik, et al.
Published: (2025)
by: Chandna, Bhavik, et al.
Published: (2025)
Gender Inclusivity Fairness Index (GIFI): A Multilevel Framework for Evaluating Gender Diversity in Large Language Models
by: Shan, Zhengyang, et al.
Published: (2025)
by: Shan, Zhengyang, et al.
Published: (2025)
LLMs Do Not See Age: Assessing Demographic Bias in Automated Systematic Review Synthesis
by: Aghaebe, Favour Yahdii, et al.
Published: (2025)
by: Aghaebe, Favour Yahdii, et al.
Published: (2025)
A Novel Method to Metigate Demographic and Expert Bias in ICD Coding with Causal Inference
by: Zhang, Bin, et al.
Published: (2024)
by: Zhang, Bin, et al.
Published: (2024)
Leveraging Prototypical Representations for Mitigating Social Bias without Demographic Information
by: Iskander, Shadi, et al.
Published: (2024)
by: Iskander, Shadi, et al.
Published: (2024)
Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability
by: Raimondi, Bianca, et al.
Published: (2025)
by: Raimondi, Bianca, et al.
Published: (2025)
Eliminating Position Bias of Language Models: A Mechanistic Approach
by: Wang, Ziqi, et al.
Published: (2024)
by: Wang, Ziqi, et al.
Published: (2024)
Trustworthy Social Bias Measurement
by: Bommasani, Rishi, et al.
Published: (2022)
by: Bommasani, Rishi, et al.
Published: (2022)
Erasing Conceptual Knowledge from Language Models
by: Gandikota, Rohit, et al.
Published: (2024)
by: Gandikota, Rohit, et al.
Published: (2024)
Web-Browsing LLMs Can Access Social Media Profiles and Infer User Demographics
by: Alizadeh, Meysam, et al.
Published: (2025)
by: Alizadeh, Meysam, et al.
Published: (2025)
Does the Prompt-based Large Language Model Recognize Students' Demographics and Introduce Bias in Essay Scoring?
by: Yang, Kaixun, et al.
Published: (2025)
by: Yang, Kaixun, et al.
Published: (2025)
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
by: Majumdar, Ayan, et al.
Published: (2025)
by: Majumdar, Ayan, et al.
Published: (2025)
The Unequal Opportunities of Large Language Models: Revealing Demographic Bias through Job Recommendations
by: Salinas, Abel, et al.
Published: (2023)
by: Salinas, Abel, et al.
Published: (2023)
Reasoning Models Can Be Effective Without Thinking
by: Ma, Wenjie, et al.
Published: (2025)
by: Ma, Wenjie, et al.
Published: (2025)
Sometimes the Model doth Preach: Quantifying Religious Bias in Open LLMs through Demographic Analysis in Asian Nations
by: Shankar, Hari, et al.
Published: (2025)
by: Shankar, Hari, et al.
Published: (2025)
Characterizing the Role of Similarity in the Property Inferences of Language Models
by: Rodriguez, Juan Diego, et al.
Published: (2024)
by: Rodriguez, Juan Diego, et al.
Published: (2024)
Don't Erase, Inform! Detecting and Contextualizing Harmful Language in Cultural Heritage Collections
by: Mastromichalakis, Orfeas Menis, et al.
Published: (2025)
by: Mastromichalakis, Orfeas Menis, et al.
Published: (2025)
Bias in News Summarization: Measures, Pitfalls and Corpora
by: Steen, Julius, et al.
Published: (2023)
by: Steen, Julius, et al.
Published: (2023)
UniErase: Towards Balanced and Precise Unlearning in Language Models
by: Yu, Miao, et al.
Published: (2025)
by: Yu, Miao, et al.
Published: (2025)
How to Evaluate Automatic Speech Recognition: Comparing Different Performance and Bias Measures
by: Patel, Tanvina, et al.
Published: (2025)
by: Patel, Tanvina, et al.
Published: (2025)
Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning
by: Ma, Chuang, et al.
Published: (2026)
by: Ma, Chuang, et al.
Published: (2026)
Steering Without Breaking: Mechanistically Informed Interventions for Discrete Diffusion Language Models
by: Zhou, Hanhan, et al.
Published: (2026)
by: Zhou, Hanhan, et al.
Published: (2026)
Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text
by: Islam, Tunazzina
Published: (2026)
by: Islam, Tunazzina
Published: (2026)
Large Language Models Share Representations of Latent Grammatical Concepts Across Typologically Diverse Languages
by: Brinkmann, Jannik, et al.
Published: (2025)
by: Brinkmann, Jannik, et al.
Published: (2025)
In-context Learning Generalizes, But Not Always Robustly: The Case of Syntax
by: Mueller, Aaron, et al.
Published: (2023)
by: Mueller, Aaron, et al.
Published: (2023)
Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases
by: Gao, Bufan, et al.
Published: (2025)
by: Gao, Bufan, et al.
Published: (2025)
Similar Items
-
MABR: Multilayer Adversarial Bias Removal Without Prior Bias Knowledge
by: Yin, Maxwell J., et al.
Published: (2024) -
Obscured but Not Erased: Evaluating Nationality Bias in LLMs via Name-Based Bias Benchmarks
by: Pelosio, Giulio, et al.
Published: (2025) -
Erasing Without Remembering: Implicit Knowledge Forgetting in Large Language Models
by: Wang, Huazheng, et al.
Published: (2025) -
In-Context Learning Without Copying
by: Sahin, Kerem, et al.
Published: (2025) -
RedacBench: Can AI Erase Your Secrets?
by: Jeon, Hyunjun, et al.
Published: (2026)