Efficient Adversarial Training in LLMs with Continuous Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Xhonneux, Sophie, Sordoni, Alessandro, Günnemann, Stephan, Gidel, Gauthier, Schwinn, Leo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
by: Dobre, David, et al.
Published: (2025)
by: Dobre, David, et al.
Published: (2025)
LLM-Safety Evaluations Lack Robustness
by: Beyer, Tim, et al.
Published: (2025)
by: Beyer, Tim, et al.
Published: (2025)
Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space
by: Schwinn, Leo, et al.
Published: (2024)
by: Schwinn, Leo, et al.
Published: (2024)
In-Context Learning Can Re-learn Forbidden Tasks
by: Xhonneux, Sophie, et al.
Published: (2024)
by: Xhonneux, Sophie, et al.
Published: (2024)
Closing the Distribution Gap in Adversarial Training for LLMs
by: Hu, Chengzhi, et al.
Published: (2026)
by: Hu, Chengzhi, et al.
Published: (2026)
Adversarial Alignment for LLMs Requires Simpler, Reproducible, and More Measurable Objectives
by: Schwinn, Leo, et al.
Published: (2025)
by: Schwinn, Leo, et al.
Published: (2025)
Adversarial Attacks on Graph Neural Networks via Meta Learning
by: Zügner, Daniel, et al.
Published: (2019)
by: Zügner, Daniel, et al.
Published: (2019)
Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs
by: Scholten, Yan, et al.
Published: (2025)
by: Scholten, Yan, et al.
Published: (2025)
Provable Adversarial Robustness for Group Equivariant Tasks: Graphs, Point Clouds, Molecules, and More
by: Schuchardt, Jan, et al.
Published: (2023)
by: Schuchardt, Jan, et al.
Published: (2023)
Provable Robustness of (Graph) Neural Networks Against Data Poisoning and Backdoor Attacks
by: Gosch, Lukas, et al.
Published: (2024)
by: Gosch, Lukas, et al.
Published: (2024)
Provable Robustness against Backdoor Attacks via the Primal-Dual Perspective on Differential Privacy
by: Saxena, Aman, et al.
Published: (2026)
by: Saxena, Aman, et al.
Published: (2026)
Fast Proxies for LLM Robustness Evaluation
by: Beyer, Tim, et al.
Published: (2025)
by: Beyer, Tim, et al.
Published: (2025)
Understanding and Improving Continuous Adversarial Training for LLMs via In-context Learning Theory
by: Fu, Shaopeng, et al.
Published: (2026)
by: Fu, Shaopeng, et al.
Published: (2026)
Generalization Properties of Adversarial Training for $\ell_0$-Bounded Adversarial Attacks
by: Delgosha, Payam, et al.
Published: (2024)
by: Delgosha, Payam, et al.
Published: (2024)
Short-length Adversarial Training Helps LLMs Defend Long-length Jailbreak Attacks: Theoretical and Empirical Evidence
by: Fu, Shaopeng, et al.
Published: (2025)
by: Fu, Shaopeng, et al.
Published: (2025)
Exact Certification of (Graph) Neural Networks Against Label Poisoning
by: Sabanayagam, Mahalakshmi, et al.
Published: (2024)
by: Sabanayagam, Mahalakshmi, et al.
Published: (2024)
Unified Mechanism-Specific Amplification by Subsampling and Group Privacy Amplification
by: Schuchardt, Jan, et al.
Published: (2024)
by: Schuchardt, Jan, et al.
Published: (2024)
Low-Cost Hard-Label Adversarial Attack with Theoretical Foundations
by: Liu, Jun, et al.
Published: (2026)
by: Liu, Jun, et al.
Published: (2026)
Enhancing Adversarial Attacks via Parameter Adaptive Adversarial Attack
by: Jin, Zhibo, et al.
Published: (2024)
by: Jin, Zhibo, et al.
Published: (2024)
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs
by: Zizzo, Giulio, et al.
Published: (2025)
by: Zizzo, Giulio, et al.
Published: (2025)
Simple and Efficient Partial Graph Adversarial Attack: A New Perspective
by: Zhu, Guanghui, et al.
Published: (2023)
by: Zhu, Guanghui, et al.
Published: (2023)
BruSLeAttack: A Query-Efficient Score-Based Black-Box Sparse Adversarial Attack
by: Vo, Viet Quoc, et al.
Published: (2024)
by: Vo, Viet Quoc, et al.
Published: (2024)
Privacy Amplification by Structured Subsampling for Deep Differentially Private Time Series Forecasting
by: Schuchardt, Jan, et al.
Published: (2025)
by: Schuchardt, Jan, et al.
Published: (2025)
Robustness Against Adversarial Attacks via Learning Confined Adversarial Polytopes
by: Hamidi, Shayan Mohajer, et al.
Published: (2024)
by: Hamidi, Shayan Mohajer, et al.
Published: (2024)
Calibration Attacks: A Comprehensive Study of Adversarial Attacks on Model Confidence
by: Obadinma, Stephen, et al.
Published: (2024)
by: Obadinma, Stephen, et al.
Published: (2024)
Attack by Unlearning: Unlearning-Induced Adversarial Attacks on Graph Neural Networks
by: Zhang, Jiahao, et al.
Published: (2026)
by: Zhang, Jiahao, et al.
Published: (2026)
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
by: Brown, Hannah, et al.
Published: (2024)
by: Brown, Hannah, et al.
Published: (2024)
Adversarial Contrastive Learning for LLM Quantization Attacks
by: Song, Dinghong, et al.
Published: (2026)
by: Song, Dinghong, et al.
Published: (2026)
Temporal Analysis of Adversarial Attacks in Federated Learning
by: Mapakshi, Rohit, et al.
Published: (2025)
by: Mapakshi, Rohit, et al.
Published: (2025)
Disttack: Graph Adversarial Attacks Toward Distributed GNN Training
by: Zhang, Yuxiang, et al.
Published: (2024)
by: Zhang, Yuxiang, et al.
Published: (2024)
Test-time Adversarial Defense with Opposite Adversarial Path and High Attack Time Cost
by: Yeh, Cheng-Han, et al.
Published: (2024)
by: Yeh, Cheng-Han, et al.
Published: (2024)
Sampling-aware Adversarial Attacks Against Large Language Models
by: Beyer, Tim, et al.
Published: (2025)
by: Beyer, Tim, et al.
Published: (2025)
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
by: Panfilov, Alexander, et al.
Published: (2026)
by: Panfilov, Alexander, et al.
Published: (2026)
Adversarial Inception Backdoor Attacks against Reinforcement Learning
by: Rathbun, Ethan, et al.
Published: (2024)
by: Rathbun, Ethan, et al.
Published: (2024)
Stealthy Adversarial Attacks on Stochastic Multi-Armed Bandits
by: Wang, Zhiwei, et al.
Published: (2024)
by: Wang, Zhiwei, et al.
Published: (2024)
Evaluating Adversarial Attacks on Federated Learning for Temperature Forecasting
by: Chichifoi, Karina, et al.
Published: (2025)
by: Chichifoi, Karina, et al.
Published: (2025)
Adversarial Attacks on Locally Private Graph Neural Networks
by: Varun, Matta, et al.
Published: (2026)
by: Varun, Matta, et al.
Published: (2026)
The Relationship Between Network Similarity and Transferability of Adversarial Attacks
by: Klause, Gerrit, et al.
Published: (2025)
by: Klause, Gerrit, et al.
Published: (2025)
Differentiable Adversarial Attacks for Marked Temporal Point Processes
by: Chakraborty, Pritish, et al.
Published: (2025)
by: Chakraborty, Pritish, et al.
Published: (2025)
Asymmetric Bias in Text-to-Image Generation with Adversarial Attacks
by: Shahgir, Haz Sameen, et al.
Published: (2023)
by: Shahgir, Haz Sameen, et al.
Published: (2023)
Similar Items
-
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
by: Dobre, David, et al.
Published: (2025) -
LLM-Safety Evaluations Lack Robustness
by: Beyer, Tim, et al.
Published: (2025) -
Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space
by: Schwinn, Leo, et al.
Published: (2024) -
In-Context Learning Can Re-learn Forbidden Tasks
by: Xhonneux, Sophie, et al.
Published: (2024) -
Closing the Distribution Gap in Adversarial Training for LLMs
by: Hu, Chengzhi, et al.
Published: (2026)