Alert-ME: An Explainability-Driven Defense Against Adversarial Examples in Transformer-Based Text Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Sabir, Bushra, Gao, Yansong, Abuadbba, Alsharif, Babar, M. Ali |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Generative Adversarial Attack for Multilingual Text Classifiers
by: Roth, Tom, et al.
Published: (2024)
by: Roth, Tom, et al.
Published: (2024)
Token-Modification Adversarial Attacks for Natural Language Processing: A Survey
by: Roth, Tom, et al.
Published: (2021)
by: Roth, Tom, et al.
Published: (2021)
Adversarial Attacks Against Automated Fact-Checking: A Survey
by: Liu, Fanzhen, et al.
Published: (2025)
by: Liu, Fanzhen, et al.
Published: (2025)
On Adversarial Examples for Text Classification by Perturbing Latent Representations
by: Sooksatra, Korn, et al.
Published: (2024)
by: Sooksatra, Korn, et al.
Published: (2024)
DDPT: Diffusion-Driven Prompt Tuning for Large Language Model Code Generation
by: Li, Jinyang, et al.
Published: (2025)
by: Li, Jinyang, et al.
Published: (2025)
Adversarial Text Purification: A Large Language Model Approach for Defense
by: Moraffah, Raha, et al.
Published: (2024)
by: Moraffah, Raha, et al.
Published: (2024)
H-FLTN: A Privacy-Preserving Hierarchical Framework for Electric Vehicle Spatio-Temporal Charge Prediction
by: Marlin, Robert, et al.
Published: (2025)
by: Marlin, Robert, et al.
Published: (2025)
MaskPure: Improving Defense Against Text Adversaries with Stochastic Purification
by: Gietz, Harrison, et al.
Published: (2024)
by: Gietz, Harrison, et al.
Published: (2024)
AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks
by: Siska, Charlotte, et al.
Published: (2025)
by: Siska, Charlotte, et al.
Published: (2025)
An Embarrassingly Simple Defense Against LLM Abliteration Attacks
by: Shairah, Harethah Abu, et al.
Published: (2025)
by: Shairah, Harethah Abu, et al.
Published: (2025)
Enhancing Hyperspace Analogue to Language (HAL) Representations via Attention-Based Pooling for Text Classification
by: Sakour, Ali, et al.
Published: (2026)
by: Sakour, Ali, et al.
Published: (2026)
From Solitary Directives to Interactive Encouragement! LLM Secure Code Generation by Natural Language Prompting
by: Liu, Shigang, et al.
Published: (2024)
by: Liu, Shigang, et al.
Published: (2024)
Explainable Transformer-Based Email Phishing Classification with Adversarial Robustness
by: P, Sajad U
Published: (2025)
by: P, Sajad U
Published: (2025)
DOGe: Defensive Output Generation for LLM Protection Against Knowledge Distillation
by: Li, Pingzhi, et al.
Published: (2025)
by: Li, Pingzhi, et al.
Published: (2025)
Analyzing the Impact of Adversarial Examples on Explainable Machine Learning
by: Devabhakthini, Prathyusha, et al.
Published: (2023)
by: Devabhakthini, Prathyusha, et al.
Published: (2023)
TaeBench: Improving Quality of Toxic Adversarial Examples
by: Zhu, Xuan, et al.
Published: (2024)
by: Zhu, Xuan, et al.
Published: (2024)
In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation
by: Chen, Shiqi, et al.
Published: (2024)
by: Chen, Shiqi, et al.
Published: (2024)
DeepiSign-G: Generic Watermark to Stamp Hidden DNN Parameters for Self-contained Tracking
by: Abuadbba, Alsharif, et al.
Published: (2024)
by: Abuadbba, Alsharif, et al.
Published: (2024)
Adversarial Training for Defense Against Label Poisoning Attacks
by: Bal, Melis Ilayda, et al.
Published: (2025)
by: Bal, Melis Ilayda, et al.
Published: (2025)
Adversarial Lens: Exploiting Attention Layers to Generate Adversarial Examples for Evaluation
by: Dhole, Kaustubh
Published: (2025)
by: Dhole, Kaustubh
Published: (2025)
Revisiting the Role of Label Smoothing in Enhanced Text Sentiment Classification
by: Gao, Yijie, et al.
Published: (2023)
by: Gao, Yijie, et al.
Published: (2023)
A Constraint-Enforcing Reward for Adversarial Attacks on Text Classifiers
by: Roth, Tom, et al.
Published: (2024)
by: Roth, Tom, et al.
Published: (2024)
Defensive Dual Masking for Robust Adversarial Defense
by: Yang, Wangli, et al.
Published: (2024)
by: Yang, Wangli, et al.
Published: (2024)
Classification of Hope in Textual Data using Transformer-Based Models
by: Ijezue, Chukwuebuka Fortunate, et al.
Published: (2025)
by: Ijezue, Chukwuebuka Fortunate, et al.
Published: (2025)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
by: Yi, Sibo, et al.
Published: (2024)
by: Yi, Sibo, et al.
Published: (2024)
LLM-Guided Semantic Bootstrapping for Interpretable Text Classification with Tsetlin Machines
by: Gao, Jiechao, et al.
Published: (2026)
by: Gao, Jiechao, et al.
Published: (2026)
AI-Generated Text Detection and Classification Based on BERT Deep Learning Algorithm
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Token Masking Improves Transformer-Based Text Classification
by: Xu, Xianglong, et al.
Published: (2025)
by: Xu, Xianglong, et al.
Published: (2025)
Towards LLM-guided Causal Explainability for Black-box Text Classifiers
by: Bhattacharjee, Amrita, et al.
Published: (2023)
by: Bhattacharjee, Amrita, et al.
Published: (2023)
Group-Adaptive Adversarial Learning for Robust Fake News Detection Against Malicious Comments
by: Tong, Zhao, et al.
Published: (2025)
by: Tong, Zhao, et al.
Published: (2025)
One Example Shown, Many Concepts Known! Counterexample-Driven Conceptual Reasoning in Mathematical LLMs
by: Li, Yinghui, et al.
Published: (2025)
by: Li, Yinghui, et al.
Published: (2025)
Evaluating Large Language Models for Health-Related Text Classification Tasks with Public Social Media Data
by: Guo, Yuting, et al.
Published: (2024)
by: Guo, Yuting, et al.
Published: (2024)
Countermeasures Against Adversarial Examples in Radio Signal Classification
by: Zhang, Lu, et al.
Published: (2024)
by: Zhang, Lu, et al.
Published: (2024)
UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
by: Lin, Huawei, et al.
Published: (2025)
by: Lin, Huawei, et al.
Published: (2025)
Adversarial Attacks on AI-Generated Text Detection Models: A Token Probability-Based Approach Using Embeddings
by: Kadhim, Ahmed K., et al.
Published: (2025)
by: Kadhim, Ahmed K., et al.
Published: (2025)
From Text to Graph: Leveraging Graph Neural Networks for Enhanced Explainability in NLP
by: Yáñez-Romero, Fabio, et al.
Published: (2025)
by: Yáñez-Romero, Fabio, et al.
Published: (2025)
Deep Adversarial Defense Against Multilevel-Lp Attacks
by: Wang, Ren, et al.
Published: (2024)
by: Wang, Ren, et al.
Published: (2024)
Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples
by: Gao, Chengqian, et al.
Published: (2025)
by: Gao, Chengqian, et al.
Published: (2025)
Can Multitask Learning Enhance Model Explainability?
by: Najjar, Hiba, et al.
Published: (2025)
by: Najjar, Hiba, et al.
Published: (2025)
CrisisKAN: Knowledge-infused and Explainable Multimodal Attention Network for Crisis Event Classification
by: Gupta, Shubham, et al.
Published: (2024)
by: Gupta, Shubham, et al.
Published: (2024)
Similar Items
-
A Generative Adversarial Attack for Multilingual Text Classifiers
by: Roth, Tom, et al.
Published: (2024) -
Token-Modification Adversarial Attacks for Natural Language Processing: A Survey
by: Roth, Tom, et al.
Published: (2021) -
Adversarial Attacks Against Automated Fact-Checking: A Survey
by: Liu, Fanzhen, et al.
Published: (2025) -
On Adversarial Examples for Text Classification by Perturbing Latent Representations
by: Sooksatra, Korn, et al.
Published: (2024) -
DDPT: Diffusion-Driven Prompt Tuning for Large Language Model Code Generation
by: Li, Jinyang, et al.
Published: (2025)