Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Zhanhui, Liu, Jie, Dong, Zhichen, Liu, Jiaheng, Yang, Chao, Ouyang, Wanli, Qiao, Yu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!