Understanding and Improving Continuous Adversarial Training for LLMs via In-context Learning Theory
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Shaopeng, Wang, Di |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Short-length Adversarial Training Helps LLMs Defend Long-length Jailbreak Attacks: Theoretical and Empirical Evidence
by: Fu, Shaopeng, et al.
Published: (2025)
by: Fu, Shaopeng, et al.
Published: (2025)
Efficient Adversarial Training in LLMs with Continuous Attacks
by: Xhonneux, Sophie, et al.
Published: (2024)
by: Xhonneux, Sophie, et al.
Published: (2024)
Pre-trained Encoder Inference: Revealing Upstream Encoders In Downstream Machine Learning Services
by: Fu, Shaopeng, et al.
Published: (2024)
by: Fu, Shaopeng, et al.
Published: (2024)
Improved Generation of Adversarial Examples Against Safety-aligned LLMs
by: Li, Qizhang, et al.
Published: (2024)
by: Li, Qizhang, et al.
Published: (2024)
Understanding In-Context Learning of Linear Models in Transformers Through an Adversarial Lens
by: Anwar, Usman, et al.
Published: (2024)
by: Anwar, Usman, et al.
Published: (2024)
Mitigating Error Amplification in Fast Adversarial Training
by: Zhao, Mengnan, et al.
Published: (2026)
by: Zhao, Mengnan, et al.
Published: (2026)
Vulnerability-Aware Robust Multimodal Adversarial Training
by: Zhang, Junrui, et al.
Published: (2025)
by: Zhang, Junrui, et al.
Published: (2025)
Robustness Against Adversarial Attacks via Learning Confined Adversarial Polytopes
by: Hamidi, Shayan Mohajer, et al.
Published: (2024)
by: Hamidi, Shayan Mohajer, et al.
Published: (2024)
Improving Clean Accuracy via a Tangent-Space Perspective on Adversarial Training
by: Yi, Bongsoo, et al.
Published: (2024)
by: Yi, Bongsoo, et al.
Published: (2024)
Robustness-Congruent Adversarial Training for Secure Machine Learning Model Updates
by: Angioni, Daniele, et al.
Published: (2024)
by: Angioni, Daniele, et al.
Published: (2024)
Generalization Properties of Adversarial Training for $\ell_0$-Bounded Adversarial Attacks
by: Delgosha, Payam, et al.
Published: (2024)
by: Delgosha, Payam, et al.
Published: (2024)
On the Effectiveness of Adversarial Training on Malware Classifiers
by: Bostani, Hamid, et al.
Published: (2024)
by: Bostani, Hamid, et al.
Published: (2024)
Understanding Deep Learning defenses Against Adversarial Examples Through Visualizations for Dynamic Risk Assessment
by: Echeberria-Barrio, Xabier, et al.
Published: (2024)
by: Echeberria-Barrio, Xabier, et al.
Published: (2024)
Adversarial Suffix Filtering: a Defense Pipeline for LLMs
by: Khachaturov, David, et al.
Published: (2025)
by: Khachaturov, David, et al.
Published: (2025)
Adversarial Machine Unlearning
by: Di, Zonglin, et al.
Published: (2024)
by: Di, Zonglin, et al.
Published: (2024)
FedBAP: Backdoor Defense via Benign Adversarial Perturbation in Federated Learning
by: Yan, Xinhai, et al.
Published: (2025)
by: Yan, Xinhai, et al.
Published: (2025)
On the Privacy Risk of In-context Learning
by: Duan, Haonan, et al.
Published: (2024)
by: Duan, Haonan, et al.
Published: (2024)
Steering Beyond the Support: Adversarial Training on Unsupervised Jailbroken Activation Simulation
by: Chen, Luoyu, et al.
Published: (2026)
by: Chen, Luoyu, et al.
Published: (2026)
Passive Inference Attacks on Split Learning via Adversarial Regularization
by: Zhu, Xiaochen, et al.
Published: (2023)
by: Zhu, Xiaochen, et al.
Published: (2023)
Adversarial Attacks on Graph Neural Networks via Meta Learning
by: Zügner, Daniel, et al.
Published: (2019)
by: Zügner, Daniel, et al.
Published: (2019)
Understanding Sensitivity of Differential Attention through the Lens of Adversarial Robustness
by: Takahashi, Tsubasa, et al.
Published: (2025)
by: Takahashi, Tsubasa, et al.
Published: (2025)
Detecting Adversarial Data via Provable Adversarial Noise Amplification
by: Mumcu, Furkan, et al.
Published: (2026)
by: Mumcu, Furkan, et al.
Published: (2026)
Enhancing Adversarial Attacks via Parameter Adaptive Adversarial Attack
by: Jin, Zhibo, et al.
Published: (2024)
by: Jin, Zhibo, et al.
Published: (2024)
Learning Robust and Privacy-Preserving Representations via Information Theory
by: Zhang, Binghui, et al.
Published: (2024)
by: Zhang, Binghui, et al.
Published: (2024)
Elevating Defenses: Bridging Adversarial Training and Watermarking for Model Resilience
by: Thakkar, Janvi, et al.
Published: (2023)
by: Thakkar, Janvi, et al.
Published: (2023)
CCLab: Adversarial Testing of Learning- and Non-Learning-Based Congestion Controllers
by: Chen, Zhi, et al.
Published: (2026)
by: Chen, Zhi, et al.
Published: (2026)
Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents
by: Kumar, Priyanshu, et al.
Published: (2024)
by: Kumar, Priyanshu, et al.
Published: (2024)
Using Graph Theory for Improving Machine Learning-based Detection of Cyber Attacks
by: Zonneveld, Giacomo, et al.
Published: (2024)
by: Zonneveld, Giacomo, et al.
Published: (2024)
DeMem: Privacy-Enhanced Robust Adversarial Learning via De-Memorization
by: Luo, Xiaoyu, et al.
Published: (2024)
by: Luo, Xiaoyu, et al.
Published: (2024)
AdvSGM: Differentially Private Graph Learning via Adversarial Skip-gram Model
by: Zhang, Sen, et al.
Published: (2025)
by: Zhang, Sen, et al.
Published: (2025)
Adaptive Meta-learning-based Adversarial Training for Robust Automatic Modulation Classification
by: Bamdad, Amirmohammad, et al.
Published: (2025)
by: Bamdad, Amirmohammad, et al.
Published: (2025)
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs
by: Zizzo, Giulio, et al.
Published: (2025)
by: Zizzo, Giulio, et al.
Published: (2025)
Intrinsic Fingerprint of LLMs: Continue Training is NOT All You Need to Steal A Model!
by: Yoon, Do-hyeon, et al.
Published: (2025)
by: Yoon, Do-hyeon, et al.
Published: (2025)
Trident: Improving Malware Detection with LLMs and Behavioral Features
by: Saul, Rebecca, et al.
Published: (2026)
by: Saul, Rebecca, et al.
Published: (2026)
Effective Universal Unrestricted Adversarial Attacks using a MOE Approach
by: Baia, A. E., et al.
Published: (2021)
by: Baia, A. E., et al.
Published: (2021)
MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment
by: Halloran, John
Published: (2025)
by: Halloran, John
Published: (2025)
Adversarial Contrastive Learning for LLM Quantization Attacks
by: Song, Dinghong, et al.
Published: (2026)
by: Song, Dinghong, et al.
Published: (2026)
Temporal Analysis of Adversarial Attacks in Federated Learning
by: Mapakshi, Rohit, et al.
Published: (2025)
by: Mapakshi, Rohit, et al.
Published: (2025)
Introducing Adaptive Continuous Adversarial Training (ACAT) to Enhance ML Robustness
by: elShehaby, Mohamed, et al.
Published: (2024)
by: elShehaby, Mohamed, et al.
Published: (2024)
How to Enhance Downstream Adversarial Robustness (almost) without Touching the Pre-Trained Foundation Model?
by: Liu, Meiqi, et al.
Published: (2025)
by: Liu, Meiqi, et al.
Published: (2025)
Similar Items
-
Short-length Adversarial Training Helps LLMs Defend Long-length Jailbreak Attacks: Theoretical and Empirical Evidence
by: Fu, Shaopeng, et al.
Published: (2025) -
Efficient Adversarial Training in LLMs with Continuous Attacks
by: Xhonneux, Sophie, et al.
Published: (2024) -
Pre-trained Encoder Inference: Revealing Upstream Encoders In Downstream Machine Learning Services
by: Fu, Shaopeng, et al.
Published: (2024) -
Improved Generation of Adversarial Examples Against Safety-aligned LLMs
by: Li, Qizhang, et al.
Published: (2024) -
Understanding In-Context Learning of Linear Models in Transformers Through an Adversarial Lens
by: Anwar, Usman, et al.
Published: (2024)