Attention Shift: Steering AI Away from Unsafe Content

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Garg, Shivank, Tiwari, Manyana
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929529499418624
author Garg, Shivank
Tiwari, Manyana
author_facet Garg, Shivank
Tiwari, Manyana
contents This study investigates the generation of unsafe or harmful content in state-of-the-art generative models, focusing on methods for restricting such generations. We introduce a novel training-free approach using attention reweighing to remove unsafe concepts without additional training during inference. We compare our method against existing ablation methods, evaluating the performance on both, direct and adversarial jailbreak prompts, using qualitative and quantitative metrics. We hypothesize potential reasons for the observed results and discuss the limitations and broader implications of content restriction.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04447
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Attention Shift: Steering AI Away from Unsafe Content
Garg, Shivank
Tiwari, Manyana
Computer Vision and Pattern Recognition
Cryptography and Security
Machine Learning
This study investigates the generation of unsafe or harmful content in state-of-the-art generative models, focusing on methods for restricting such generations. We introduce a novel training-free approach using attention reweighing to remove unsafe concepts without additional training during inference. We compare our method against existing ablation methods, evaluating the performance on both, direct and adversarial jailbreak prompts, using qualitative and quantitative metrics. We hypothesize potential reasons for the observed results and discuss the limitations and broader implications of content restriction.
title Attention Shift: Steering AI Away from Unsafe Content
topic Computer Vision and Pattern Recognition
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2410.04447