Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Yi, Xin, Li, Yue, Shi, Dongsheng, Wang, Linlin, Wang, Xiaoling, He, Liang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From static to adaptive: immune memory-based jailbreak detection for large language models
by: Leng, Jun, et al.
Published: (2025)
by: Leng, Jun, et al.
Published: (2025)
TeleAI-Safety: A comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluations
by: Chen, Xiuyuan, et al.
Published: (2025)
by: Chen, Xiuyuan, et al.
Published: (2025)
Nonideality-aware training makes memristive networks more robust to adversarial attacks
by: Joksas, Dovydas, et al.
Published: (2024)
by: Joksas, Dovydas, et al.
Published: (2024)
On the use of neurosymbolic AI for defending against cyber attacks
by: Grov, Gudmund, et al.
Published: (2024)
by: Grov, Gudmund, et al.
Published: (2024)
Improving behavior based authentication against adversarial attack using XAI
by: Qin, Dong, et al.
Published: (2024)
by: Qin, Dong, et al.
Published: (2024)
Recursive language models for jailbreak detection: a procedural defense for tool-augmented agents
by: Shavit, Doron
Published: (2026)
by: Shavit, Doron
Published: (2026)
SHLIME: Foiling adversarial attacks fooling SHAP and LIME
by: Chauhan, Sam, et al.
Published: (2025)
by: Chauhan, Sam, et al.
Published: (2025)
Problem space structural adversarial attacks for Network Intrusion Detection Systems based on Graph Neural Networks
by: Venturi, Andrea, et al.
Published: (2024)
by: Venturi, Andrea, et al.
Published: (2024)
Precision-Varying Prediction (PVP): Robustifying ASR systems against adversarial attacks
by: Pizarro, Matías, et al.
Published: (2026)
by: Pizarro, Matías, et al.
Published: (2026)
Revisiting the attacker's knowledge in inference attacks against Searchable Symmetric Encryption
by: Damie, Marc, et al.
Published: (2025)
by: Damie, Marc, et al.
Published: (2025)
Spoofing attack augmentation: can differently-trained attack models improve generalisation?
by: Ge, Wanying, et al.
Published: (2023)
by: Ge, Wanying, et al.
Published: (2023)
Assessing biomedical knowledge robustness in large language models by query-efficient sampling attacks
by: Xian, R. Patrick, et al.
Published: (2024)
by: Xian, R. Patrick, et al.
Published: (2024)
AGMark: Attention-Guided Dynamic Watermarking for Large Vision-Language Models
by: Li, Yue, et al.
Published: (2026)
by: Li, Yue, et al.
Published: (2026)
PMANet: Malicious URL detection via post-trained language model guided multi-level feature attention network
by: Liu, Ruitong, et al.
Published: (2023)
by: Liu, Ruitong, et al.
Published: (2023)
RAB$^2$-DEF: Dynamic and explainable defense against adversarial attacks in Federated Learning to fair poor clients
by: Rodríguez-Barroso, Nuria, et al.
Published: (2024)
by: Rodríguez-Barroso, Nuria, et al.
Published: (2024)
Exploring the limits of strong membership inference attacks on large language models
by: Hayes, Jamie, et al.
Published: (2025)
by: Hayes, Jamie, et al.
Published: (2025)
Enhancing the Robustness of QMIX against State-adversarial Attacks
by: Guo, Weiran, et al.
Published: (2023)
by: Guo, Weiran, et al.
Published: (2023)
A traffic analysis attack against Introduction Protocol and Onion Services
by: Constantinides, Nicolas
Published: (2026)
by: Constantinides, Nicolas
Published: (2026)
DLP: towards active defense against backdoor attacks with decoupled learning process
by: Ying, Zonghao, et al.
Published: (2024)
by: Ying, Zonghao, et al.
Published: (2024)
Malacopula: adversarial automatic speaker verification attacks using a neural-based generalised Hammerstein model
by: Todisco, Massimiliano, et al.
Published: (2024)
by: Todisco, Massimiliano, et al.
Published: (2024)
Trainwreck: A damaging adversarial attack on image classifiers
by: Zahálka, Jan
Published: (2023)
by: Zahálka, Jan
Published: (2023)
Learning diverse attacks on large language models for robust red-teaming and safety tuning
by: Lee, Seanie, et al.
Published: (2024)
by: Lee, Seanie, et al.
Published: (2024)
Prompt Injection attack against LLM-integrated Applications
by: Liu, Yi, et al.
Published: (2023)
by: Liu, Yi, et al.
Published: (2023)
Switching multiplicative watermark design against covert attacks
by: Gallo, Alexander J., et al.
Published: (2025)
by: Gallo, Alexander J., et al.
Published: (2025)
Adversarial attacks against Modern Vision-Language Models
by: La Torre, Alejandro Paredes
Published: (2026)
by: La Torre, Alejandro Paredes
Published: (2026)
Secret extraction attacks against obfuscated IQP circuits
by: Gross, David, et al.
Published: (2023)
by: Gross, David, et al.
Published: (2023)
Correlation inference attacks against machine learning models
by: Creţu, Ana-Maria, et al.
Published: (2021)
by: Creţu, Ana-Maria, et al.
Published: (2021)
FreeTalk:A plug-and-play and black-box defense against speech synthesis attacks
by: Pu, Yuwen, et al.
Published: (2025)
by: Pu, Yuwen, et al.
Published: (2025)
Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction
by: Chen, Yulin, et al.
Published: (2025)
by: Chen, Yulin, et al.
Published: (2025)
On the existence of consistent adversarial attacks in high-dimensional linear classification
by: Vilucchio, Matteo, et al.
Published: (2025)
by: Vilucchio, Matteo, et al.
Published: (2025)
Analysis of the vulnerability of machine learning regression models to adversarial attacks using data from 5G wireless networks
by: Legashev, Leonid, et al.
Published: (2025)
by: Legashev, Leonid, et al.
Published: (2025)
Security for adversarial wiretap channels
by: Hänggi, Esther, et al.
Published: (2024)
by: Hänggi, Esther, et al.
Published: (2024)
Noisy Neighbors: Efficient membership inference attacks against LLMs
by: Galli, Filippo, et al.
Published: (2024)
by: Galli, Filippo, et al.
Published: (2024)
Limits of privacy amplification against non-signalling memory attacks
by: Arnon, Rotem, et al.
Published: (2012)
by: Arnon, Rotem, et al.
Published: (2012)
Quantum forgery attacks against OTR structures based on Simon's algorithm
by: Liu, Wenjie, et al.
Published: (2023)
by: Liu, Wenjie, et al.
Published: (2023)
Requiem for a drone: a machine-learning based framework for stealthy attacks against unmanned autonomous vehicles
by: Kim, Kyo Hyun, et al.
Published: (2024)
by: Kim, Kyo Hyun, et al.
Published: (2024)
Acoustic Interference: A New Paradigm Weaponizing Acoustic Latent Semantic for Universal Jailbreak against Large Audio Language Models
by: Wang, Yanyun, et al.
Published: (2026)
by: Wang, Yanyun, et al.
Published: (2026)
Cross-site scripting adversarial attacks based on deep reinforcement learning: Evaluation and extension study
by: Pasini, Samuele, et al.
Published: (2025)
by: Pasini, Samuele, et al.
Published: (2025)
Power side-channel leakage localization through adversarial training of deep neural networks
by: Gammell, Jimmy, et al.
Published: (2024)
by: Gammell, Jimmy, et al.
Published: (2024)
Threat analysis and adversarial model for Smart Grids
by: Ríos, Javier Sande, et al.
Published: (2024)
by: Ríos, Javier Sande, et al.
Published: (2024)
Similar Items
-
From static to adaptive: immune memory-based jailbreak detection for large language models
by: Leng, Jun, et al.
Published: (2025) -
TeleAI-Safety: A comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluations
by: Chen, Xiuyuan, et al.
Published: (2025) -
Nonideality-aware training makes memristive networks more robust to adversarial attacks
by: Joksas, Dovydas, et al.
Published: (2024) -
On the use of neurosymbolic AI for defending against cyber attacks
by: Grov, Gudmund, et al.
Published: (2024) -
Improving behavior based authentication against adversarial attack using XAI
by: Qin, Dong, et al.
Published: (2024)