Saved in:
| Main Authors: | , |
|---|---|
| Format: | Recurso digital |
| Language: | |
| Published: |
Zenodo
2026
|
| Online Access: | https://doi.org/10.5281/zenodo.19833311 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Table of Contents:
- <div> <div class="standard-markdown grid-cols-1 grid [&_>_*]:min-w-0 gap-3 standard-markdown"> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]"> </p> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">Title: When Explanations Lie: XAI Reliability Under Adversarial Obfuscation in Phishing Detection</p> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">Description:</p> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">This paper presents a systematic empirical audit of Explainable Artificial Intelligence (XAI) reliability for phishing and spam email detection under realistic adversarial text obfuscation, specifically the character substitution and token fragmentation techniques routinely employed by real attackers.</p> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">The central contribution is the introduction of Contradiction Rate (CR), a formally defined novel metric that detects a failure mode not captured by any existing XAI reliability measure: when the aggregate signed attribution of the top-k explanation features collectively opposes the model's own predicted class, rendering the explanation actively misleading to an analyst. Unlike faithfulness scores, which measure confidence drop on feature removal, or Jaccard agreement, which measures feature set overlap without regard to sign, CR directly quantifies the proportion of instances in which explanations point in the direction opposite to the model's decision. A deployable CR-based Explanation Health Monitor with Healthy, Degraded, and Critical threshold states is derived from the clean-data CR baseline.</p> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">Using two datasets, SpamAssassin (5,809 emails) and MeAJOR (108,685 emails, published 2026), and three model classes (TF-IDF Logistic Regression, TF-IDF Linear SVM, and DistilBERT), the paper evaluates SHAP and LIME explanation reliability under four adversarial transforms: leetspeak character substitution, homoglyph Unicode replacement, whitespace token splitting, and an unmodified clean baseline.</p> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">Four principal findings are reported. First, under leetspeak, SHAP-LIME inter-explainer agreement drops from 0.393 to 0.261 while CR increases 21-fold from 0.020 to 0.425, meaning that in 42.5% of obfuscated instances the explanation collectively argues against the model's own classification. Second, global explanation stability is paradoxically inverse to model accuracy: a low-dimensional 8-feature Random Forest achieves Spearman rank stability of 0.986 across random seeds versus just 0.069 for a high-dimensional TF-IDF model with far superior classification accuracy, a phenomenon the authors term the Accuracy-Stability Paradox and the XAI selection trap. Third, explanation reliability degrades independently of classification accuracy under obfuscation, demonstrated using the Explanation Stability Score as a normalised cross-metric comparability device. Fourth, DistilBERT, despite achieving 95.3% clean accuracy, collapses to 49.2% mean accuracy under leetspeak, a 48.4% drop compared to 31.5% for classical TF-IDF models, showing that model sophistication does not confer adversarial robustness under character-level obfuscation.</p> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">The findings carry direct implications for Security Operations Centre deployments of XAI-assisted threat triage and for compliance with EU AI Act transparency requirements. All experiments are reproducible using publicly available datasets and standard open-source libraries.</p> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">Keywords: Explainable Artificial Intelligence, XAI reliability, Contradiction Rate, phishing detection, adversarial obfuscation, SHAP, LIME, DistilBERT, cybersecurity, SOC, explanation failure, leetspeak, homoglyph</p> </div> </div>