Safety Alignment of Large Language Models via Contrasting Safe and Harmful Distributions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Xiaoyun, Zhao, Zhengyue, Shi, Wenxuan, Xu, Kaidi, Huang, Di, Hu, Xing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!