Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yi, Xin, Li, Yue, Shi, Dongsheng, Wang, Linlin, Wang, Xiaoling, He, Liang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!