Can Global XAI Methods Reveal Injected Bias in LLMs? SHAP vs Rule Extraction vs RuleSHAP
Fuente:
arXiv
Guardado en:
| Autor principal: | Sovrano, Francesco |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Neuron-Anchored Rule Extraction for Large Language Models via Contrastive Hierarchical Ablation
por: Sovrano, Francesco, et al.
Publicado: (2026)
por: Sovrano, Francesco, et al.
Publicado: (2026)
PolySHAP: Extending KernelSHAP with Interaction-Informed Polynomial Regression
por: Fumagalli, Fabian, et al.
Publicado: (2026)
por: Fumagalli, Fabian, et al.
Publicado: (2026)
Towards Piece-by-Piece Explanations for Chess Positions with SHAP
por: Spinnato, Francesco
Publicado: (2025)
por: Spinnato, Francesco
Publicado: (2025)
Interaction Tensor SHAP
por: Hasegawa, Hiroki, et al.
Publicado: (2025)
por: Hasegawa, Hiroki, et al.
Publicado: (2025)
Towards trustable SHAP scores
por: Letoffe, Olivier, et al.
Publicado: (2024)
por: Letoffe, Olivier, et al.
Publicado: (2024)
CS-SHAP: Extending SHAP to Cyclic-Spectral Domain for Better Interpretability of Intelligent Fault Diagnosis
por: Chen, Qian, et al.
Publicado: (2025)
por: Chen, Qian, et al.
Publicado: (2025)
A Perspective on Explainable Artificial Intelligence Methods: SHAP and LIME
por: Salih, Ahmed, et al.
Publicado: (2023)
por: Salih, Ahmed, et al.
Publicado: (2023)
ContextualSHAP : Enhancing SHAP Explanations Through Contextual Language Generation
por: Dwiyanti, Latifa, et al.
Publicado: (2025)
por: Dwiyanti, Latifa, et al.
Publicado: (2025)
Explanation Multiplicity in SHAP: Characterization and Assessment
por: Hwang, Hyunseung, et al.
Publicado: (2026)
por: Hwang, Hyunseung, et al.
Publicado: (2026)
SHAP-based Explanations are Sensitive to Feature Representation
por: Hwang, Hyunseung, et al.
Publicado: (2025)
por: Hwang, Hyunseung, et al.
Publicado: (2025)
From SHAP Scores to Feature Importance Scores
por: Letoffe, Olivier, et al.
Publicado: (2024)
por: Letoffe, Olivier, et al.
Publicado: (2024)
On the Tractability of SHAP Explanations under Markovian Distributions
por: Marzouk, Reda, et al.
Publicado: (2024)
por: Marzouk, Reda, et al.
Publicado: (2024)
Fooling SHAP with Output Shuffling Attacks
por: Yuan, Jun, et al.
Publicado: (2024)
por: Yuan, Jun, et al.
Publicado: (2024)
Can LLMs Follow Simple Rules?
por: Mu, Norman, et al.
Publicado: (2023)
por: Mu, Norman, et al.
Publicado: (2023)
HyperSHAP: Shapley Values and Interactions for Explaining Hyperparameter Optimization
por: Wever, Marcel, et al.
Publicado: (2025)
por: Wever, Marcel, et al.
Publicado: (2025)
How to safely discard features based on aggregate SHAP values
por: Bhattacharjee, Robi, et al.
Publicado: (2025)
por: Bhattacharjee, Robi, et al.
Publicado: (2025)
SHAP scores fail pervasively even when Lipschitz succeeds
por: Letoffe, Olivier, et al.
Publicado: (2024)
por: Letoffe, Olivier, et al.
Publicado: (2024)
A Polynomial-Time Axiomatic Alternative to SHAP for Feature Attribution
por: Hiraki, Kazuhiro, et al.
Publicado: (2026)
por: Hiraki, Kazuhiro, et al.
Publicado: (2026)
CORTEX: A Cost-Sensitive Rule and Tree Extraction Method
por: Kopanja, Marija, et al.
Publicado: (2025)
por: Kopanja, Marija, et al.
Publicado: (2025)
CQD-SHAP: Explainable Complex Query Answering via Shapley Values
por: Abbasi, Parsa, et al.
Publicado: (2025)
por: Abbasi, Parsa, et al.
Publicado: (2025)
Enhancing SHAP Explainability for Diagnostic and Prognostic ML Models in Alzheimer Disease
por: Guillén, Pablo, et al.
Publicado: (2026)
por: Guillén, Pablo, et al.
Publicado: (2026)
KernelSHAP-IQ: Weighted Least-Square Optimization for Shapley Interactions
por: Fumagalli, Fabian, et al.
Publicado: (2024)
por: Fumagalli, Fabian, et al.
Publicado: (2024)
DeltaSHAP: Explaining Prediction Evolutions in Online Patient Monitoring with Shapley Values
por: Kim, Changhun, et al.
Publicado: (2025)
por: Kim, Changhun, et al.
Publicado: (2025)
Shaping Up SHAP: Enhancing Stability through Layer-Wise Neighbor Selection
por: Kelodjou, Gwladys, et al.
Publicado: (2023)
por: Kelodjou, Gwladys, et al.
Publicado: (2023)
Causal SHAP: Feature Attribution with Dependency Awareness through Causal Discovery
por: Ng, Woon Yee, et al.
Publicado: (2025)
por: Ng, Woon Yee, et al.
Publicado: (2025)
Enhancing RL Generalizability in Robotics through SHAP Analysis of Algorithms and Hyperparameters
por: Kong, Lingxiao, et al.
Publicado: (2026)
por: Kong, Lingxiao, et al.
Publicado: (2026)
CNN-TFT explained by SHAP with multi-head attention weights for time series forecasting
por: Stefenon, Stefano F., et al.
Publicado: (2025)
por: Stefenon, Stefano F., et al.
Publicado: (2025)
TN-SHAP-G: Graph-Structured Tensor Network Surrogates for Shapley Values and Interactions
por: Heidari, Farzaneh, et al.
Publicado: (2026)
por: Heidari, Farzaneh, et al.
Publicado: (2026)
Choose Your Explanation: A Comparison of SHAP and GradCAM in Human Activity Recognition
por: Tempel, Felix, et al.
Publicado: (2024)
por: Tempel, Felix, et al.
Publicado: (2024)
FairSHAP: Preprocessing for Fairness Through Attribution-Based Data Augmentation
por: Zhu, Lin, et al.
Publicado: (2025)
por: Zhu, Lin, et al.
Publicado: (2025)
Enabling Regional Explainability by Automatic and Model-agnostic Rule Extraction
por: Chen, Yu, et al.
Publicado: (2024)
por: Chen, Yu, et al.
Publicado: (2024)
Automatic Extraction of Linguistic Description from Fuzzy Rule Base
por: Siminski, Krzysztof, et al.
Publicado: (2024)
por: Siminski, Krzysztof, et al.
Publicado: (2024)
InjectTST: A Transformer Method of Injecting Global Information into Independent Channels for Long Time Series Forecasting
por: Chi, Ce, et al.
Publicado: (2024)
por: Chi, Ce, et al.
Publicado: (2024)
Verified SHAP: Provable Bounds for Exact Shapley Values of Neural Networks
por: Boetius, David, et al.
Publicado: (2026)
por: Boetius, David, et al.
Publicado: (2026)
Interpreting Time Series Forecasts with LIME and SHAP: A Case Study on the Air Passengers Dataset
por: Shukla, Manish
Publicado: (2025)
por: Shukla, Manish
Publicado: (2025)
SHAP Distance: An Explainability-Aware Metric for Evaluating the Semantic Fidelity of Synthetic Tabular Data
por: Yu, Ke, et al.
Publicado: (2025)
por: Yu, Ke, et al.
Publicado: (2025)
Detecting Cybersecurity Threats by Integrating Explainable AI with SHAP Interpretability and Strategic Data Sampling
por: Srisumrith, Norrakith, et al.
Publicado: (2026)
por: Srisumrith, Norrakith, et al.
Publicado: (2026)
Tabular Foundation Models Can Learn Association Rules
por: Karabulut, Erkan, et al.
Publicado: (2026)
por: Karabulut, Erkan, et al.
Publicado: (2026)
In-Context Symbolic Regression for Robustness-Improved Kolmogorov-Arnold Networks
por: Sovrano, Francesco, et al.
Publicado: (2026)
por: Sovrano, Francesco, et al.
Publicado: (2026)
Injecting Structured Biomedical Knowledge into Language Models: Continual Pretraining vs. GraphRAG
por: Klila, Jaafer, et al.
Publicado: (2026)
por: Klila, Jaafer, et al.
Publicado: (2026)
Ejemplares similares
-
Neuron-Anchored Rule Extraction for Large Language Models via Contrastive Hierarchical Ablation
por: Sovrano, Francesco, et al.
Publicado: (2026) -
PolySHAP: Extending KernelSHAP with Interaction-Informed Polynomial Regression
por: Fumagalli, Fabian, et al.
Publicado: (2026) -
Towards Piece-by-Piece Explanations for Chess Positions with SHAP
por: Spinnato, Francesco
Publicado: (2025) -
Interaction Tensor SHAP
por: Hasegawa, Hiroki, et al.
Publicado: (2025) -
Towards trustable SHAP scores
por: Letoffe, Olivier, et al.
Publicado: (2024)