Interpreting Adversarial Attacks and Defences using Architectures with Enhanced Interpretability
Fuente:
arXiv
Saved in:
| Main Authors: | Rao, Akshay G, Lakshminarayanan, Chandrashekhar, Rajkumar, Arun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Interpretability-Guided Test-Time Adversarial Defense
by: Kulkarni, Akshay, et al.
Published: (2024)
by: Kulkarni, Akshay, et al.
Published: (2024)
Enhancing Adversarial Attacks via Parameter Adaptive Adversarial Attack
by: Jin, Zhibo, et al.
Published: (2024)
by: Jin, Zhibo, et al.
Published: (2024)
Why You Should Not Trust Interpretations in Machine Learning: Adversarial Attacks on Partial Dependence Plots
by: Xin, Xi, et al.
Published: (2024)
by: Xin, Xi, et al.
Published: (2024)
Adaptive Randomized Smoothing: Certified Adversarial Robustness for Multi-Step Defences
by: Lyu, Saiyue, et al.
Published: (2024)
by: Lyu, Saiyue, et al.
Published: (2024)
Effective Universal Unrestricted Adversarial Attacks using a MOE Approach
by: Baia, A. E., et al.
Published: (2021)
by: Baia, A. E., et al.
Published: (2021)
AUTOLYCUS: Exploiting Explainable AI (XAI) for Model Extraction Attacks against Interpretable Models
by: Oksuz, Abdullah Caglar, et al.
Published: (2023)
by: Oksuz, Abdullah Caglar, et al.
Published: (2023)
Robustness Against Adversarial Attacks via Learning Confined Adversarial Polytopes
by: Hamidi, Shayan Mohajer, et al.
Published: (2024)
by: Hamidi, Shayan Mohajer, et al.
Published: (2024)
Generalization Properties of Adversarial Training for $\ell_0$-Bounded Adversarial Attacks
by: Delgosha, Payam, et al.
Published: (2024)
by: Delgosha, Payam, et al.
Published: (2024)
Calibration Attacks: A Comprehensive Study of Adversarial Attacks on Model Confidence
by: Obadinma, Stephen, et al.
Published: (2024)
by: Obadinma, Stephen, et al.
Published: (2024)
Attack by Unlearning: Unlearning-Induced Adversarial Attacks on Graph Neural Networks
by: Zhang, Jiahao, et al.
Published: (2026)
by: Zhang, Jiahao, et al.
Published: (2026)
BB-Patch: BlackBox Adversarial Patch-Attack using Zeroth-Order Optimization
by: Kumar, Satyadwyoom, et al.
Published: (2024)
by: Kumar, Satyadwyoom, et al.
Published: (2024)
Laplace Transform Interpretation of Differential Privacy
by: Chourasia, Rishav, et al.
Published: (2024)
by: Chourasia, Rishav, et al.
Published: (2024)
Temporal Analysis of Adversarial Attacks in Federated Learning
by: Mapakshi, Rohit, et al.
Published: (2025)
by: Mapakshi, Rohit, et al.
Published: (2025)
Adversarial Contrastive Learning for LLM Quantization Attacks
by: Song, Dinghong, et al.
Published: (2026)
by: Song, Dinghong, et al.
Published: (2026)
Efficient Adversarial Training in LLMs with Continuous Attacks
by: Xhonneux, Sophie, et al.
Published: (2024)
by: Xhonneux, Sophie, et al.
Published: (2024)
Autonomous Network Defence using Reinforcement Learning
by: Foley, Myles, et al.
Published: (2024)
by: Foley, Myles, et al.
Published: (2024)
Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs
by: Davies, Xander, et al.
Published: (2025)
by: Davies, Xander, et al.
Published: (2025)
Entity-based Reinforcement Learning for Autonomous Cyber Defence
by: Thompson, Isaac Symes, et al.
Published: (2024)
by: Thompson, Isaac Symes, et al.
Published: (2024)
Test-time Adversarial Defense with Opposite Adversarial Path and High Attack Time Cost
by: Yeh, Cheng-Han, et al.
Published: (2024)
by: Yeh, Cheng-Han, et al.
Published: (2024)
Using Retriever Augmented Large Language Models for Attack Graph Generation
by: Prapty, Renascence Tarafder, et al.
Published: (2024)
by: Prapty, Renascence Tarafder, et al.
Published: (2024)
Evaluating Adversarial Attacks on Federated Learning for Temperature Forecasting
by: Chichifoi, Karina, et al.
Published: (2025)
by: Chichifoi, Karina, et al.
Published: (2025)
The Relationship Between Network Similarity and Transferability of Adversarial Attacks
by: Klause, Gerrit, et al.
Published: (2025)
by: Klause, Gerrit, et al.
Published: (2025)
Differentiable Adversarial Attacks for Marked Temporal Point Processes
by: Chakraborty, Pritish, et al.
Published: (2025)
by: Chakraborty, Pritish, et al.
Published: (2025)
Adversarial Attacks on Locally Private Graph Neural Networks
by: Varun, Matta, et al.
Published: (2026)
by: Varun, Matta, et al.
Published: (2026)
Adversarial Inception Backdoor Attacks against Reinforcement Learning
by: Rathbun, Ethan, et al.
Published: (2024)
by: Rathbun, Ethan, et al.
Published: (2024)
Stealthy Adversarial Attacks on Stochastic Multi-Armed Bandits
by: Wang, Zhiwei, et al.
Published: (2024)
by: Wang, Zhiwei, et al.
Published: (2024)
Asymmetric Bias in Text-to-Image Generation with Adversarial Attacks
by: Shahgir, Haz Sameen, et al.
Published: (2023)
by: Shahgir, Haz Sameen, et al.
Published: (2023)
Colliding with Adversaries at ECML-PKDD 2025 Adversarial Attack Competition 1st Prize Solution
by: Stefanopoulos, Dimitris, et al.
Published: (2025)
by: Stefanopoulos, Dimitris, et al.
Published: (2025)
Constrained Adaptive Attack: Effective Adversarial Attack Against Deep Neural Networks for Tabular Data
by: Simonetto, Thibault, et al.
Published: (2024)
by: Simonetto, Thibault, et al.
Published: (2024)
Real-World Adversarial Attacks on RF-Based Drone Detectors
by: Gazit, Omer, et al.
Published: (2025)
by: Gazit, Omer, et al.
Published: (2025)
A Comprehensive Analysis of Adversarial Attacks against Spam Filters
by: Hotoğlu, Esra, et al.
Published: (2025)
by: Hotoğlu, Esra, et al.
Published: (2025)
Moshi Moshi? A Model Selection Hijacking Adversarial Attack
by: Petrucci, Riccardo, et al.
Published: (2025)
by: Petrucci, Riccardo, et al.
Published: (2025)
Passive Inference Attacks on Split Learning via Adversarial Regularization
by: Zhu, Xiaochen, et al.
Published: (2023)
by: Zhu, Xiaochen, et al.
Published: (2023)
Low-Cost Hard-Label Adversarial Attack with Theoretical Foundations
by: Liu, Jun, et al.
Published: (2026)
by: Liu, Jun, et al.
Published: (2026)
Explainability-Based Adversarial Attack on Graphs Through Edge Perturbation
by: Chanda, Dibaloke, et al.
Published: (2023)
by: Chanda, Dibaloke, et al.
Published: (2023)
Adversarial Attacks on Graph Neural Networks via Meta Learning
by: Zügner, Daniel, et al.
Published: (2019)
by: Zügner, Daniel, et al.
Published: (2019)
BruSLeAttack: A Query-Efficient Score-Based Black-Box Sparse Adversarial Attack
by: Vo, Viet Quoc, et al.
Published: (2024)
by: Vo, Viet Quoc, et al.
Published: (2024)
Adversarially-Aware Architecture Design for Robust Medical AI Systems
by: Gerhart, Alyssa, et al.
Published: (2025)
by: Gerhart, Alyssa, et al.
Published: (2025)
FABLE: A Localized, Targeted Adversarial Attack on Weather Forecasting Models
by: Deng, Yue, et al.
Published: (2025)
by: Deng, Yue, et al.
Published: (2025)
Intriguing Properties of Adversarial ML Attacks in the Problem Space [Extended Version]
by: Cortellazzi, Jacopo, et al.
Published: (2019)
by: Cortellazzi, Jacopo, et al.
Published: (2019)
Similar Items
-
Interpretability-Guided Test-Time Adversarial Defense
by: Kulkarni, Akshay, et al.
Published: (2024) -
Enhancing Adversarial Attacks via Parameter Adaptive Adversarial Attack
by: Jin, Zhibo, et al.
Published: (2024) -
Why You Should Not Trust Interpretations in Machine Learning: Adversarial Attacks on Partial Dependence Plots
by: Xin, Xi, et al.
Published: (2024) -
Adaptive Randomized Smoothing: Certified Adversarial Robustness for Multi-Step Defences
by: Lyu, Saiyue, et al.
Published: (2024) -
Effective Universal Unrestricted Adversarial Attacks using a MOE Approach
by: Baia, A. E., et al.
Published: (2021)