Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yuxi, Liu, Yi, Li, Yuekang, Shi, Ling, Deng, Gelei, Chen, Shengquan, Wang, Kailong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917416409235456
author Li, Yuxi
Liu, Yi
Li, Yuekang
Shi, Ling
Deng, Gelei
Chen, Shengquan
Wang, Kailong
author_facet Li, Yuxi
Liu, Yi
Li, Yuekang
Shi, Ling
Deng, Gelei
Chen, Shengquan
Wang, Kailong
contents Large language models (LLMs) have revolutionized various applications, making robust safety alignment essential to prevent harmful outputs. Current safety alignment techniques, however, harbor inherent vulnerabilities due to their reliance on logit suppression. In this work, we identify critical logit-level vulnerabilities by introducing Semantic-sensitive Alignment and Generation (SSAG), a method designed to systematically manipulate output-layer logits without altering model parameters. Experiments on five popular LLMs show that SSAG exposes harmful responses with a 95% success rate while reducing response time by 86%. VulMine also demonstrates superior attack efficacy, achieving an average ASR of up to 77% against strong defensive mechanisms. These findings reveal crucial weaknesses in existing alignment methods, highlighting an urgent need for improved vulnerability detection and robust safety alignment strategies. Our code is available on github.
format Preprint
id arxiv_https___arxiv_org_abs_2405_13068
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
Li, Yuxi
Liu, Yi
Li, Yuekang
Shi, Ling
Deng, Gelei
Chen, Shengquan
Wang, Kailong
Cryptography and Security
Artificial Intelligence
Machine Learning
Large language models (LLMs) have revolutionized various applications, making robust safety alignment essential to prevent harmful outputs. Current safety alignment techniques, however, harbor inherent vulnerabilities due to their reliance on logit suppression. In this work, we identify critical logit-level vulnerabilities by introducing Semantic-sensitive Alignment and Generation (SSAG), a method designed to systematically manipulate output-layer logits without altering model parameters. Experiments on five popular LLMs show that SSAG exposes harmful responses with a 95% success rate while reducing response time by 86%. VulMine also demonstrates superior attack efficacy, achieving an average ASR of up to 77% against strong defensive mechanisms. These findings reveal crucial weaknesses in existing alignment methods, highlighting an urgent need for improved vulnerability detection and robust safety alignment strategies. Our code is available on github.
title Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2405.13068