Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bhardwaj, Rishabh, Anh, Do Duc, Poria, Soujanya
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!