Defending Against Unforeseen Failure Modes with Latent Adversarial Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Casper, Stephen, Schulze, Lennart, Patel, Oam, Hadfield-Menell, Dylan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Prompt Injection as Role Confusion
von: Ye, Charles, et al.
Veröffentlicht: (2026)
von: Ye, Charles, et al.
Veröffentlicht: (2026)
Calibrated Adversarial Sampling: Multi-Armed Bandit-Guided Generalization Against Unforeseen Attacks
von: Wang, Rui, et al.
Veröffentlicht: (2025)
von: Wang, Rui, et al.
Veröffentlicht: (2025)
Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks
von: Halloran, John T., et al.
Veröffentlicht: (2026)
von: Halloran, John T., et al.
Veröffentlicht: (2026)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
von: Cao, Bochuan, et al.
Veröffentlicht: (2023)
von: Cao, Bochuan, et al.
Veröffentlicht: (2023)
A White-Box Adversarial Attack Against a Digital Twin
von: Patterson, Wilson, et al.
Veröffentlicht: (2022)
von: Patterson, Wilson, et al.
Veröffentlicht: (2022)
MoCo-EA: Exploiting Adversarial Mode Connectivity for Efficient Evolutionary Attacks
von: Kim, Hyo Seo, et al.
Veröffentlicht: (2026)
von: Kim, Hyo Seo, et al.
Veröffentlicht: (2026)
Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States
von: Chia, Xin Wei, et al.
Veröffentlicht: (2025)
von: Chia, Xin Wei, et al.
Veröffentlicht: (2025)
Information Theoretic Adversarial Training of Large Language Models
von: Zhang, Yiwei, et al.
Veröffentlicht: (2026)
von: Zhang, Yiwei, et al.
Veröffentlicht: (2026)
One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries
von: Zloczower, Itay, et al.
Veröffentlicht: (2026)
von: Zloczower, Itay, et al.
Veröffentlicht: (2026)
Your Agent Can Defend Itself against Backdoor Attacks
von: Changjiang, Li, et al.
Veröffentlicht: (2025)
von: Changjiang, Li, et al.
Veröffentlicht: (2025)
Efficient but Vulnerable: Benchmarking and Defending LLM Batch Prompting Attack
von: Yue, Murong, et al.
Veröffentlicht: (2025)
von: Yue, Murong, et al.
Veröffentlicht: (2025)
Defending against Adversarial Malware Attacks on ML-based Android Malware Detection Systems
von: He, Ping, et al.
Veröffentlicht: (2025)
von: He, Ping, et al.
Veröffentlicht: (2025)
Disttack: Graph Adversarial Attacks Toward Distributed GNN Training
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2024)
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2024)
Embedding Hidden Adversarial Capabilities in Pre-Trained Diffusion Models
von: Beerens, Lucas, et al.
Veröffentlicht: (2025)
von: Beerens, Lucas, et al.
Veröffentlicht: (2025)
The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
von: Jia, Feiran, et al.
Veröffentlicht: (2024)
von: Jia, Feiran, et al.
Veröffentlicht: (2024)
Learning to Defend by Attacking (and Vice-Versa): Transfer of Learning in Cybersecurity Games
von: Malloy, Tailia, et al.
Veröffentlicht: (2023)
von: Malloy, Tailia, et al.
Veröffentlicht: (2023)
Multi-class Classifier based Failure Prediction with Artificial and Anonymous Training for Data Privacy
von: Das, Dibakar, et al.
Veröffentlicht: (2022)
von: Das, Dibakar, et al.
Veröffentlicht: (2022)
NPAT Null-Space Projected Adversarial Training Towards Zero Deterioration
von: Hu, Hanyi, et al.
Veröffentlicht: (2024)
von: Hu, Hanyi, et al.
Veröffentlicht: (2024)
Defending Against Poisoning Attacks in Federated Learning with Blockchain
von: Dong, Nanqing, et al.
Veröffentlicht: (2023)
von: Dong, Nanqing, et al.
Veröffentlicht: (2023)
Optimal Defender Strategies for CAGE-2 using Causal Modeling and Tree Search
von: Hammar, Kim, et al.
Veröffentlicht: (2024)
von: Hammar, Kim, et al.
Veröffentlicht: (2024)
Defending the Edge: Representative-Attention Defense against Backdoor Attacks in Federated Learning
von: Obioma, Chibueze Peace, et al.
Veröffentlicht: (2025)
von: Obioma, Chibueze Peace, et al.
Veröffentlicht: (2025)
Attackers Strike Back? Not Anymore -- An Ensemble of RL Defenders Awakens for APT Detection
von: Benabderrahmane, Sidahmed, et al.
Veröffentlicht: (2025)
von: Benabderrahmane, Sidahmed, et al.
Veröffentlicht: (2025)
Improving Clean Accuracy via a Tangent-Space Perspective on Adversarial Training
von: Yi, Bongsoo, et al.
Veröffentlicht: (2024)
von: Yi, Bongsoo, et al.
Veröffentlicht: (2024)
Unveiling the Backdoor Mechanism Hidden Behind Catastrophic Overfitting in Fast Adversarial Training
von: Zhao, Mengnan, et al.
Veröffentlicht: (2026)
von: Zhao, Mengnan, et al.
Veröffentlicht: (2026)
Generalist++: A Meta-learning Framework for Mitigating Trade-off in Adversarial Training
von: Wang, Yisen, et al.
Veröffentlicht: (2025)
von: Wang, Yisen, et al.
Veröffentlicht: (2025)
Have You Poisoned My Data? Defending Neural Networks against Data Poisoning
von: De Gaspari, Fabio, et al.
Veröffentlicht: (2024)
von: De Gaspari, Fabio, et al.
Veröffentlicht: (2024)
Demystifying the Mythos or Disrupting Bugonomics? From Zero-Day Asymmetry to Defender Remediation Throughput
von: Pesoli, Alfredo, et al.
Veröffentlicht: (2026)
von: Pesoli, Alfredo, et al.
Veröffentlicht: (2026)
Trapping Attacker in Dilemma: Examining Internal Correlations and External Influences of Trigger for Defending GNN Backdoors
von: Yang, Fan, et al.
Veröffentlicht: (2026)
von: Yang, Fan, et al.
Veröffentlicht: (2026)
SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection
von: Chen, Shuhao, et al.
Veröffentlicht: (2026)
von: Chen, Shuhao, et al.
Veröffentlicht: (2026)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
von: Liu, Fan, et al.
Veröffentlicht: (2024)
von: Liu, Fan, et al.
Veröffentlicht: (2024)
Why Does Differential Privacy with Large Epsilon Defend Against Practical Membership Inference Attacks?
von: Lowy, Andrew, et al.
Veröffentlicht: (2024)
von: Lowy, Andrew, et al.
Veröffentlicht: (2024)
PubDef: Defending Against Transfer Attacks From Public Models
von: Sitawarin, Chawin, et al.
Veröffentlicht: (2023)
von: Sitawarin, Chawin, et al.
Veröffentlicht: (2023)
Can Adversarial Code Comments Fool AI Security Reviewers -- Large-Scale Empirical Study of Comment-Based Attacks and Defenses Against LLM Code Analysis
von: Thornton, Scott
Veröffentlicht: (2026)
von: Thornton, Scott
Veröffentlicht: (2026)
Fight Back Against Jailbreaking via Prompt Adversarial Tuning
von: Mo, Yichuan, et al.
Veröffentlicht: (2024)
von: Mo, Yichuan, et al.
Veröffentlicht: (2024)
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
von: Liang, Buyun, et al.
Veröffentlicht: (2026)
von: Liang, Buyun, et al.
Veröffentlicht: (2026)
Feature Selection via GANs (GANFS): Enhancing Machine Learning Models for DDoS Mitigation
von: Patel, Harsh
Veröffentlicht: (2025)
von: Patel, Harsh
Veröffentlicht: (2025)
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
von: Che, Zora, et al.
Veröffentlicht: (2025)
von: Che, Zora, et al.
Veröffentlicht: (2025)
Vision Transformer with Adversarial Indicator Token against Adversarial Attacks in Radio Signal Classifications
von: Zhang, Lu, et al.
Veröffentlicht: (2025)
von: Zhang, Lu, et al.
Veröffentlicht: (2025)
Attacks and Defenses Against LLM Fingerprinting
von: Kurian, Kevin, et al.
Veröffentlicht: (2025)
von: Kurian, Kevin, et al.
Veröffentlicht: (2025)
Self-interpreting Adversarial Images
von: Zhang, Tingwei, et al.
Veröffentlicht: (2024)
von: Zhang, Tingwei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Prompt Injection as Role Confusion
von: Ye, Charles, et al.
Veröffentlicht: (2026) -
Calibrated Adversarial Sampling: Multi-Armed Bandit-Guided Generalization Against Unforeseen Attacks
von: Wang, Rui, et al.
Veröffentlicht: (2025) -
Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks
von: Halloran, John T., et al.
Veröffentlicht: (2026) -
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
von: Cao, Bochuan, et al.
Veröffentlicht: (2023) -
A White-Box Adversarial Attack Against a Digital Twin
von: Patterson, Wilson, et al.
Veröffentlicht: (2022)