Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods
Fuente:
arXiv
Saved in:
| Main Authors: | Dhaini, Mahdi, Erdogan, Ege, Feldhus, Nils, Kasneci, Gjergji |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EvalxNLP: A Framework for Benchmarking Post-Hoc Explainability Methods on NLP Models
by: Dhaini, Mahdi, et al.
Published: (2025)
by: Dhaini, Mahdi, et al.
Published: (2025)
When Explainability Meets Privacy: An Investigation at the Intersection of Post-hoc Explainability and Differential Privacy in the Context of Natural Language Processing
by: Dhaini, Mahdi, et al.
Published: (2025)
by: Dhaini, Mahdi, et al.
Published: (2025)
Can LLM-Generated Textual Explanations Enhance Model Classification Performance? An Empirical Study
by: Dhaini, Mahdi, et al.
Published: (2025)
by: Dhaini, Mahdi, et al.
Published: (2025)
Adoption of Explainable Natural Language Processing: Perspectives from Industry and Academia on Practices and Challenges
by: Dhaini, Mahdi, et al.
Published: (2025)
by: Dhaini, Mahdi, et al.
Published: (2025)
Position: Uncertainty Quantification Needs Reassessment for Large-language Model Agents
by: Kirchhof, Michael, et al.
Published: (2025)
by: Kirchhof, Michael, et al.
Published: (2025)
Emergent Abilities in Large Language Models: A Survey
by: Berti, Leonardo, et al.
Published: (2025)
by: Berti, Leonardo, et al.
Published: (2025)
Interpreting Language Models Through Concept Descriptions: A Survey
by: Feldhus, Nils, et al.
Published: (2025)
by: Feldhus, Nils, et al.
Published: (2025)
Simplifying Outcomes of Language Model Component Analyses with ELIA
by: Eidt, Aaron Louis, et al.
Published: (2026)
by: Eidt, Aaron Louis, et al.
Published: (2026)
CURE: Controlled Unlearning for Robust Embeddings -- Mitigating Conceptual Shortcuts in Pre-Trained Language Models
by: Kocak, Aysenur, et al.
Published: (2025)
by: Kocak, Aysenur, et al.
Published: (2025)
Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers
by: Leemann, Tobias, et al.
Published: (2024)
by: Leemann, Tobias, et al.
Published: (2024)
Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization
by: Sun, Jingyi, et al.
Published: (2026)
by: Sun, Jingyi, et al.
Published: (2026)
The Effect of Model Size on LLM Post-hoc Explainability via LIME
by: Heyen, Henning, et al.
Published: (2024)
by: Heyen, Henning, et al.
Published: (2024)
Enriching Tabular Data with Contextual LLM Embeddings: A Comprehensive Ablation Study for Ensemble Classifiers
by: Kasneci, Gjergji, et al.
Published: (2024)
by: Kasneci, Gjergji, et al.
Published: (2024)
Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
by: Wang, Qianli, et al.
Published: (2026)
by: Wang, Qianli, et al.
Published: (2026)
RAZOR: Sharpening Knowledge by Cutting Bias with Unsupervised Text Rewriting
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
Consolidating Rewarded Perturbations for LLM Post-Training
by: Zhang, Zheyu, et al.
Published: (2026)
by: Zhang, Zheyu, et al.
Published: (2026)
Is Crowdsourcing Breaking Your Bank? Cost-Effective Fine-Tuning of Pre-trained Language Models with Proximal Policy Optimization
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
Understanding Post-hoc Explainers: The Case of Anchors
by: Lopardo, Gianluigi, et al.
Published: (2023)
by: Lopardo, Gianluigi, et al.
Published: (2023)
Understanding Knowledge Drift in LLMs through Misinformation
by: Fastowski, Alina, et al.
Published: (2024)
by: Fastowski, Alina, et al.
Published: (2024)
LLMCheckup: Conversational Examination of Large Language Models via Interpretability Tools and Self-Explanations
by: Wang, Qianli, et al.
Published: (2024)
by: Wang, Qianli, et al.
Published: (2024)
Assessing the Real-World Utility of Explainable AI for Arousal Diagnostics: An Application-Grounded User Study
by: Kraft, Stefan, et al.
Published: (2025)
by: Kraft, Stefan, et al.
Published: (2025)
From Confidence to Collapse in LLM Factual Robustness
by: Fastowski, Alina, et al.
Published: (2025)
by: Fastowski, Alina, et al.
Published: (2025)
Grokking in the Wild: Data Augmentation for Real-World Multi-Hop Reasoning with Transformers
by: Abramov, Roman, et al.
Published: (2025)
by: Abramov, Roman, et al.
Published: (2025)
Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework
by: Kopf, Laura, et al.
Published: (2025)
by: Kopf, Laura, et al.
Published: (2025)
Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models
by: Zhao, Dachuan, et al.
Published: (2025)
by: Zhao, Dachuan, et al.
Published: (2025)
In Defence of Post-hoc Explainability
by: Oh, Nick
Published: (2024)
by: Oh, Nick
Published: (2024)
Mitigating Extrinsic Gender Bias for Bangla Classification Tasks
by: Joy, Sajib Kumar Saha, et al.
Published: (2024)
by: Joy, Sajib Kumar Saha, et al.
Published: (2024)
Where Paths Split: Localized, Calibrated Control of Moral Reasoning in Large Language Models
by: Yuan, Chenchen, et al.
Published: (2026)
by: Yuan, Chenchen, et al.
Published: (2026)
Facts Fade Fast: Evaluating Memorization of Outdated Medical Knowledge in Large Language Models
by: Vladika, Juraj, et al.
Published: (2025)
by: Vladika, Juraj, et al.
Published: (2025)
PXGen: A Post-hoc Explainable Method for Generative Models
by: Huang, Yen-Lung, et al.
Published: (2025)
by: Huang, Yen-Lung, et al.
Published: (2025)
I Prefer not to Say: Protecting User Consent in Models with Optional Personal Data
by: Leemann, Tobias, et al.
Published: (2022)
by: Leemann, Tobias, et al.
Published: (2022)
MoESD: Mixture of Experts Stable Diffusion to Mitigate Gender Bias
by: Wang, Guorun, et al.
Published: (2024)
by: Wang, Guorun, et al.
Published: (2024)
TLOB: A Novel Transformer Model with Dual Attention for Price Trend Prediction with Limit Order Book Data
by: Berti, Leonardo, et al.
Published: (2025)
by: Berti, Leonardo, et al.
Published: (2025)
Evaluating Gender Bias Transfer between Pre-trained and Prompt-Adapted Language Models
by: Mackraz, Natalie, et al.
Published: (2024)
by: Mackraz, Natalie, et al.
Published: (2024)
Investigating Annotator Bias in Large Language Models for Hate Speech Detection
by: Das, Amit, et al.
Published: (2024)
by: Das, Amit, et al.
Published: (2024)
Benchmarking Gender and Political Bias in Large Language Models
by: Yang, Jinrui, et al.
Published: (2025)
by: Yang, Jinrui, et al.
Published: (2025)
Doubling Your Data in Minutes: Ultra-fast Tabular Data Generation via LLM-Induced Dependency Graphs
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
Active Tabular Augmentation via Policy-Guided Diffusion Inpainting
by: Zhang, Zheyu, et al.
Published: (2026)
by: Zhang, Zheyu, et al.
Published: (2026)
What's in a Name? Auditing Large Language Models for Race and Gender Bias
by: Salinas, Alejandro, et al.
Published: (2024)
by: Salinas, Alejandro, et al.
Published: (2024)
Bridging the Bosphorus: Advancing Turkish Large Language Models through Strategies for Low-Resource Language Adaptation and Benchmarking
by: Acikgoz, Emre Can, et al.
Published: (2024)
by: Acikgoz, Emre Can, et al.
Published: (2024)
Similar Items
-
EvalxNLP: A Framework for Benchmarking Post-Hoc Explainability Methods on NLP Models
by: Dhaini, Mahdi, et al.
Published: (2025) -
When Explainability Meets Privacy: An Investigation at the Intersection of Post-hoc Explainability and Differential Privacy in the Context of Natural Language Processing
by: Dhaini, Mahdi, et al.
Published: (2025) -
Can LLM-Generated Textual Explanations Enhance Model Classification Performance? An Empirical Study
by: Dhaini, Mahdi, et al.
Published: (2025) -
Adoption of Explainable Natural Language Processing: Perspectives from Industry and Academia on Practices and Challenges
by: Dhaini, Mahdi, et al.
Published: (2025) -
Position: Uncertainty Quantification Needs Reassessment for Large-language Model Agents
by: Kirchhof, Michael, et al.
Published: (2025)