Adversarial attacks and defenses in explainable artificial intelligence: A survey
Fuente:
arXiv
Saved in:
| Main Authors: | Baniecki, Hubert, Biecek, Przemyslaw |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Interpreting CLIP with Hierarchical Sparse Autoencoders
by: Zaigrajew, Vladimir, et al.
Published: (2025)
by: Zaigrajew, Vladimir, et al.
Published: (2025)
Birds look like cars: Adversarial analysis of intrinsically interpretable deep learning
by: Baniecki, Hubert, et al.
Published: (2025)
by: Baniecki, Hubert, et al.
Published: (2025)
Shadow defense against gradient inversion attack in federated learning
by: Jiang, Le, et al.
Published: (2025)
by: Jiang, Le, et al.
Published: (2025)
SwordBench: Evaluating Orthogonality of Steering Image Representations
by: Zaigrajew, Vladimir, et al.
Published: (2026)
by: Zaigrajew, Vladimir, et al.
Published: (2026)
Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf Interactions
by: Baniecki, Hubert, et al.
Published: (2025)
by: Baniecki, Hubert, et al.
Published: (2025)
On the explainable properties of 1-Lipschitz Neural Networks: An Optimal Transport Perspective
by: Serrurier, Mathieu, et al.
Published: (2022)
by: Serrurier, Mathieu, et al.
Published: (2022)
Exploring the Adversarial Frontier: Quantifying Robustness via Adversarial Hypervolume
by: Guo, Ping, et al.
Published: (2024)
by: Guo, Ping, et al.
Published: (2024)
Red-Teaming Segment Anything Model
by: Jankowski, Krzysztof, et al.
Published: (2024)
by: Jankowski, Krzysztof, et al.
Published: (2024)
Adversarial Detection by Approximation of Ensemble Boundary
by: Windeatt, T.
Published: (2022)
by: Windeatt, T.
Published: (2022)
A Survey on Physical Adversarial Attacks against Face Recognition Systems
by: Wang, Mingsi, et al.
Published: (2024)
by: Wang, Mingsi, et al.
Published: (2024)
MOS-Attack: A Scalable Multi-objective Adversarial Attack Framework
by: Guo, Ping, et al.
Published: (2025)
by: Guo, Ping, et al.
Published: (2025)
AdvSecureNet: A Python Toolkit for Adversarial Machine Learning
by: Catal, Melih, et al.
Published: (2024)
by: Catal, Melih, et al.
Published: (2024)
On the Importance of Backbone to the Adversarial Robustness of Object Detectors
by: Li, Xiao, et al.
Published: (2023)
by: Li, Xiao, et al.
Published: (2023)
Fingerprinting Image-to-Image Generative Adversarial Networks
by: Li, Guanlin, et al.
Published: (2021)
by: Li, Guanlin, et al.
Published: (2021)
Evaluating the Evaluators: Trust in Adversarial Robustness Tests
by: Cinà, Antonio Emanuele, et al.
Published: (2025)
by: Cinà, Antonio Emanuele, et al.
Published: (2025)
Improving the Transferability of Adversarial Attacks by an Input Transpose
by: Wan, Qing, et al.
Published: (2025)
by: Wan, Qing, et al.
Published: (2025)
GLEAN: Generative Learning for Eliminating Adversarial Noise
by: Kim, Justin Lyu, et al.
Published: (2024)
by: Kim, Justin Lyu, et al.
Published: (2024)
The Impact of Scaling Training Data on Adversarial Robustness
by: Zimmerli, Marco, et al.
Published: (2025)
by: Zimmerli, Marco, et al.
Published: (2025)
AR-GAN: Generative Adversarial Network-Based Defense Method Against Adversarial Attacks on the Traffic Sign Classification System of Autonomous Vehicles
by: Salek, M Sabbir, et al.
Published: (2023)
by: Salek, M Sabbir, et al.
Published: (2023)
Sy-FAR: Symmetry-based Fair Adversarial Robustness
by: Najjar, Haneen, et al.
Published: (2025)
by: Najjar, Haneen, et al.
Published: (2025)
Towards Understanding Dual BN In Hybrid Adversarial Training
by: Zhang, Chenshuang, et al.
Published: (2024)
by: Zhang, Chenshuang, et al.
Published: (2024)
Standard-Deviation-Inspired Regularization for Improving Adversarial Robustness
by: Fakorede, Olukorede, et al.
Published: (2024)
by: Fakorede, Olukorede, et al.
Published: (2024)
Impact of Architectural Modifications on Deep Learning Adversarial Robustness
by: Juraev, Firuz, et al.
Published: (2024)
by: Juraev, Firuz, et al.
Published: (2024)
Interpretation of Neural Networks is Susceptible to Universal Adversarial Perturbations
by: Oskouie, Haniyeh Ehsani, et al.
Published: (2022)
by: Oskouie, Haniyeh Ehsani, et al.
Published: (2022)
DIFFender: Diffusion-Based Adversarial Defense against Patch Attacks
by: Kang, Caixin, et al.
Published: (2023)
by: Kang, Caixin, et al.
Published: (2023)
Undermining Image and Text Classification Algorithms Using Adversarial Attacks
by: Lunga, Langalibalele, et al.
Published: (2024)
by: Lunga, Langalibalele, et al.
Published: (2024)
Improving Adversarial Training using Vulnerability-Aware Perturbation Budget
by: Fakorede, Olukorede, et al.
Published: (2024)
by: Fakorede, Olukorede, et al.
Published: (2024)
A Survey on the Application of Generative Adversarial Networks in Cybersecurity: Prospective, Direction and Open Research Scopes
by: Arifin, Md Mashrur, et al.
Published: (2024)
by: Arifin, Md Mashrur, et al.
Published: (2024)
Pixel Seal: Adversarial-only training for invisible image and video watermarking
by: Souček, Tomáš, et al.
Published: (2025)
by: Souček, Tomáš, et al.
Published: (2025)
GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMs
by: Basani, Advik Raj, et al.
Published: (2024)
by: Basani, Advik Raj, et al.
Published: (2024)
Towards Million-Scale Adversarial Robustness Evaluation With Stronger Individual Attacks
by: Xie, Yong, et al.
Published: (2024)
by: Xie, Yong, et al.
Published: (2024)
Position: Explain to Question not to Justify
by: Biecek, Przemyslaw, et al.
Published: (2024)
by: Biecek, Przemyslaw, et al.
Published: (2024)
Securing Visually-Aware Recommender Systems: An Adversarial Image Reconstruction and Detection Framework
by: Yin, Minglei, et al.
Published: (2023)
by: Yin, Minglei, et al.
Published: (2023)
DarkLLM: Learning Language-Driven Adversarial Attacks with Large Language Models
by: Sun, Ye, et al.
Published: (2026)
by: Sun, Ye, et al.
Published: (2026)
Reinforcement Learning Platform for Adversarial Black-box Attacks with Custom Distortion Filters
by: Sarkar, Soumyendu, et al.
Published: (2025)
by: Sarkar, Soumyendu, et al.
Published: (2025)
Real-world Adversarial Defense against Patch Attacks based on Diffusion Model
by: Wei, Xingxing, et al.
Published: (2024)
by: Wei, Xingxing, et al.
Published: (2024)
Understanding Key Point Cloud Features for Development Three-dimensional Adversarial Attacks
by: Naderi, Hanieh, et al.
Published: (2022)
by: Naderi, Hanieh, et al.
Published: (2022)
Crafting Adversarial Inputs for Large Vision-Language Models Using Black-Box Optimization
by: Guan, Jiwei, et al.
Published: (2026)
by: Guan, Jiwei, et al.
Published: (2026)
Efficient Semi-Supervised Adversarial Training via Latent Clustering-Based Data Reduction
by: Ghosh, Somrita, et al.
Published: (2025)
by: Ghosh, Somrita, et al.
Published: (2025)
Stop Reasoning! When Multimodal LLM with Chain-of-Thought Reasoning Meets Adversarial Image
by: Wang, Zefeng, et al.
Published: (2024)
by: Wang, Zefeng, et al.
Published: (2024)
Similar Items
-
Interpreting CLIP with Hierarchical Sparse Autoencoders
by: Zaigrajew, Vladimir, et al.
Published: (2025) -
Birds look like cars: Adversarial analysis of intrinsically interpretable deep learning
by: Baniecki, Hubert, et al.
Published: (2025) -
Shadow defense against gradient inversion attack in federated learning
by: Jiang, Le, et al.
Published: (2025) -
SwordBench: Evaluating Orthogonality of Steering Image Representations
by: Zaigrajew, Vladimir, et al.
Published: (2026) -
Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf Interactions
by: Baniecki, Hubert, et al.
Published: (2025)