RL-JACK: Reinforcement Learning-powered Black-box Jailbreaking Attack against LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Xuan, Nie, Yuzhou, Yan, Lu, Mao, Yunshu, Guo, Wenbo, Zhang, Xiangyu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!