Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917416409235456 |
|---|---|
| author | Li, Yuxi Liu, Yi Li, Yuekang Shi, Ling Deng, Gelei Chen, Shengquan Wang, Kailong |
| author_facet | Li, Yuxi Liu, Yi Li, Yuekang Shi, Ling Deng, Gelei Chen, Shengquan Wang, Kailong |
| contents | Large language models (LLMs) have revolutionized various applications, making robust safety alignment essential to prevent harmful outputs. Current safety alignment techniques, however, harbor inherent vulnerabilities due to their reliance on logit suppression. In this work, we identify critical logit-level vulnerabilities by introducing Semantic-sensitive Alignment and Generation (SSAG), a method designed to systematically manipulate output-layer logits without altering model parameters. Experiments on five popular LLMs show that SSAG exposes harmful responses with a 95% success rate while reducing response time by 86%. VulMine also demonstrates superior attack efficacy, achieving an average ASR of up to 77% against strong defensive mechanisms. These findings reveal crucial weaknesses in existing alignment methods, highlighting an urgent need for improved vulnerability detection and robust safety alignment strategies. Our code is available on github. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_13068 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment Li, Yuxi Liu, Yi Li, Yuekang Shi, Ling Deng, Gelei Chen, Shengquan Wang, Kailong Cryptography and Security Artificial Intelligence Machine Learning Large language models (LLMs) have revolutionized various applications, making robust safety alignment essential to prevent harmful outputs. Current safety alignment techniques, however, harbor inherent vulnerabilities due to their reliance on logit suppression. In this work, we identify critical logit-level vulnerabilities by introducing Semantic-sensitive Alignment and Generation (SSAG), a method designed to systematically manipulate output-layer logits without altering model parameters. Experiments on five popular LLMs show that SSAG exposes harmful responses with a 95% success rate while reducing response time by 86%. VulMine also demonstrates superior attack efficacy, achieving an average ASR of up to 77% against strong defensive mechanisms. These findings reveal crucial weaknesses in existing alignment methods, highlighting an urgent need for improved vulnerability detection and robust safety alignment strategies. Our code is available on github. |
| title | Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment |
| topic | Cryptography and Security Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2405.13068 |