When Interpretability Becomes a Liability: Adversarial Attacks on CBM Concept Layers
Fuente:
arXiv
Saved in:
| Main Author: | Sridhar, Aditya |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Labels Have Structure: Improving Image Classification with Hierarchy-Aware Cross-Entropy
by: Chan, April, et al.
Published: (2026)
by: Chan, April, et al.
Published: (2026)
Breaking the Illusion of Security via Interpretation: Interpretable Vision Transformer Systems under Attack
by: Abdukhamidov, Eldor, et al.
Published: (2025)
by: Abdukhamidov, Eldor, et al.
Published: (2025)
TRIGS: Trojan Identification from Gradient-based Signatures
by: Hussein, Mohamed E., et al.
Published: (2023)
by: Hussein, Mohamed E., et al.
Published: (2023)
Boosting Ray Search Procedure of Hard-label Attacks with Transfer-based Priors
by: Ma, Chen, et al.
Published: (2025)
by: Ma, Chen, et al.
Published: (2025)
A Multidisciplinary Approach to Telegram Data Analysis
by: Varbanov, Velizar, et al.
Published: (2024)
by: Varbanov, Velizar, et al.
Published: (2024)
CoCoA-Mix: Confusion-and-Confidence-Aware Mixture Model for Context Optimization
by: Hong, Dasol, et al.
Published: (2025)
by: Hong, Dasol, et al.
Published: (2025)
Localizing Adversarial Attacks To Produces More Imperceptible Noise
by: Reddy, Pavan, et al.
Published: (2025)
by: Reddy, Pavan, et al.
Published: (2025)
Riemannian-Geometric Fingerprints of Generative Models
by: Song, Hae Jin, et al.
Published: (2025)
by: Song, Hae Jin, et al.
Published: (2025)
How Can One Choose the Best CAM-Based Explainability Method for a CNN Model?
by: Costa, Daniel da Silva, et al.
Published: (2026)
by: Costa, Daniel da Silva, et al.
Published: (2026)
BadSAD: Clean-Label Backdoor Attacks against Deep Semi-Supervised Anomaly Detection
by: Cheng, He, et al.
Published: (2024)
by: Cheng, He, et al.
Published: (2024)
CAMRI Loss: Improving Recall of a Specific Class without Sacrificing Accuracy
by: Nishiyama, Daiki, et al.
Published: (2022)
by: Nishiyama, Daiki, et al.
Published: (2022)
Ternary Decision Trees with Locally-Adaptive Uncertainty Zones
by: Smits, William
Published: (2026)
by: Smits, William
Published: (2026)
A Boundary-Aware Non-parametric Granular-Ball Classifier Based on Minimum Description Length
by: Xian, Zeqiang, et al.
Published: (2026)
by: Xian, Zeqiang, et al.
Published: (2026)
RobustModelMaker: Coupling Bootstrap Stability Selection with Leakage-Safe Nested Cross-Validation for Scientific Machine Learning
by: Barnard, Amanda S
Published: (2026)
by: Barnard, Amanda S
Published: (2026)
Density-aware Sample-specific Attack
by: Wang, Qiyuan, et al.
Published: (2026)
by: Wang, Qiyuan, et al.
Published: (2026)
Spiking Neural Networks for event-based action recognition: A new task to understand their advantage
by: Vicente-Sola, Alex, et al.
Published: (2022)
by: Vicente-Sola, Alex, et al.
Published: (2022)
Signal-Based Malware Classification Using 1D CNNs
by: Wilkie, Jack, et al.
Published: (2025)
by: Wilkie, Jack, et al.
Published: (2025)
Poison in the Well: Feature Embedding Disruption in Backdoor Attacks
by: Feng, Zhou, et al.
Published: (2025)
by: Feng, Zhou, et al.
Published: (2025)
A Spatio-Temporal Deep Learning Approach For High-Resolution Gridded Monsoon Prediction
by: Borah, Parashjyoti, et al.
Published: (2026)
by: Borah, Parashjyoti, et al.
Published: (2026)
Modeling the Neonatal Brain Development Using Implicit Neural Representations
by: Bieder, Florentin, et al.
Published: (2024)
by: Bieder, Florentin, et al.
Published: (2024)
Interpreting Structured Perturbations in Image Protection Methods for Diffusion Models
by: Martin, Michael R., et al.
Published: (2025)
by: Martin, Michael R., et al.
Published: (2025)
Non-Adaptive Adversarial Face Generation
by: Kim, Sunpill, et al.
Published: (2025)
by: Kim, Sunpill, et al.
Published: (2025)
NM-Hebb: Coupling Local Hebbian Plasticity with Metric Learning for More Accurate and Interpretable CNNs
by: Miličević, Davorin, et al.
Published: (2025)
by: Miličević, Davorin, et al.
Published: (2025)
When Does Margin Clamping Affect Training Variance? Dataset-Dependent Effects in Contrastive Forward-Forward Learning
by: Steier, Joshua
Published: (2026)
by: Steier, Joshua
Published: (2026)
Breaking the Stealth-Potency Trade-off in Clean-Image Backdoors with Generative Trigger Optimization
by: Xu, Binyan, et al.
Published: (2025)
by: Xu, Binyan, et al.
Published: (2025)
Improving the Convergence Rate of Ray Search Optimization for Query-Efficient Hard-Label Attacks
by: Xu, Xinjie, et al.
Published: (2025)
by: Xu, Xinjie, et al.
Published: (2025)
Adversarial Attacks and Defenses in Fault Detection and Diagnosis: A Comprehensive Benchmark on the Tennessee Eastman Process
by: Pozdnyakov, Vitaliy, et al.
Published: (2024)
by: Pozdnyakov, Vitaliy, et al.
Published: (2024)
Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection
by: Lelle, Travis
Published: (2026)
by: Lelle, Travis
Published: (2026)
On the Development of Binary Classification Algorithm Based on Principles of Geometry and Statistical Inference
by: Srivastava, Vatsal
Published: (2025)
by: Srivastava, Vatsal
Published: (2025)
MM-Food-100K: A 100,000-Sample Multimodal Food Intelligence Dataset with Verifiable Provenance
by: Dong, Yi, et al.
Published: (2025)
by: Dong, Yi, et al.
Published: (2025)
Closing the Distribution Gap in Adversarial Training for LLMs
by: Hu, Chengzhi, et al.
Published: (2026)
by: Hu, Chengzhi, et al.
Published: (2026)
Et Tu Certifications: Robustness Certificates Yield Better Adversarial Examples
by: Cullen, Andrew C., et al.
Published: (2023)
by: Cullen, Andrew C., et al.
Published: (2023)
Generalizable and Interpretable RF Fingerprinting with Shapelet-Enhanced Large Language Models
by: Zhao, Tianya, et al.
Published: (2026)
by: Zhao, Tianya, et al.
Published: (2026)
CASE: Contrastive Activation for Saliency Estimation
by: Williamson, Dane, et al.
Published: (2025)
by: Williamson, Dane, et al.
Published: (2025)
Understanding Adversarial Transferability in Vision-Language Models for Autonomous Driving: A Cross-Architecture Analysis
by: Fernandez, David, et al.
Published: (2026)
by: Fernandez, David, et al.
Published: (2026)
One Surrogate to Fool Them All: Universal, Transferable, and Targeted Adversarial Attacks with CLIP
by: Xu, Binyan, et al.
Published: (2025)
by: Xu, Binyan, et al.
Published: (2025)
NeuroViz: Real-time Interactive Visualization of Forward and Backward Passes in Neural Network Training
by: Sharma, Tanvi, et al.
Published: (2026)
by: Sharma, Tanvi, et al.
Published: (2026)
ScaleMAP: Preserving Local Density and Neighborhood Structure in Low-Dimensional Embeddings
by: Poorna, Rajas, et al.
Published: (2026)
by: Poorna, Rajas, et al.
Published: (2026)
Semantically Guided Adversarial Testing of Vision Models Using Language Models
by: Filus, Katarzyna, et al.
Published: (2025)
by: Filus, Katarzyna, et al.
Published: (2025)
Scalable Deep Subspace Clustering Network
by: Mrabah, Nairouz, et al.
Published: (2025)
by: Mrabah, Nairouz, et al.
Published: (2025)
Similar Items
-
When Labels Have Structure: Improving Image Classification with Hierarchy-Aware Cross-Entropy
by: Chan, April, et al.
Published: (2026) -
Breaking the Illusion of Security via Interpretation: Interpretable Vision Transformer Systems under Attack
by: Abdukhamidov, Eldor, et al.
Published: (2025) -
TRIGS: Trojan Identification from Gradient-based Signatures
by: Hussein, Mohamed E., et al.
Published: (2023) -
Boosting Ray Search Procedure of Hard-label Attacks with Transfer-based Priors
by: Ma, Chen, et al.
Published: (2025) -
A Multidisciplinary Approach to Telegram Data Analysis
by: Varbanov, Velizar, et al.
Published: (2024)