XBreaking: Understanding how LLMs security alignment can be broken

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Arazzi, Marco, Kembu, Vignesh Kumar, Nocera, Antonino, P, Vinod
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917066170171392
author Arazzi, Marco
Kembu, Vignesh Kumar
Nocera, Antonino
P, Vinod
author_facet Arazzi, Marco
Kembu, Vignesh Kumar
Nocera, Antonino
P, Vinod
contents Large Language Models are fundamental actors in the modern IT landscape dominated by AI solutions. However, security threats associated with them might prevent their reliable adoption in critical application scenarios such as government organizations and medical institutions. For this reason, commercial LLMs typically undergo a sophisticated censoring mechanism to eliminate any harmful output they could possibly produce. These mechanisms maintain the integrity of LLM alignment by guaranteeing that the models respond safely and ethically. In response to this, attacks on LLMs are a significant threat to such protections, and many previous approaches have already demonstrated their effectiveness across diverse domains. Existing LLM attacks mostly adopt a generate-and-test strategy to craft malicious input. To improve the comprehension of censoring mechanisms and design a targeted attack, we propose an Explainable-AI solution that comparatively analyzes the behavior of censored and uncensored models to derive unique exploitable alignment patterns. Then, we propose XBreaking, a novel approach that exploits these unique patterns to break the security and alignment constraints of LLMs by targeted noise injection. Our thorough experimental campaign returns important insights about the censoring mechanisms and demonstrates the effectiveness and performance of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2504_21700
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle XBreaking: Understanding how LLMs security alignment can be broken
Arazzi, Marco
Kembu, Vignesh Kumar
Nocera, Antonino
P, Vinod
Cryptography and Security
Artificial Intelligence
Machine Learning
Large Language Models are fundamental actors in the modern IT landscape dominated by AI solutions. However, security threats associated with them might prevent their reliable adoption in critical application scenarios such as government organizations and medical institutions. For this reason, commercial LLMs typically undergo a sophisticated censoring mechanism to eliminate any harmful output they could possibly produce. These mechanisms maintain the integrity of LLM alignment by guaranteeing that the models respond safely and ethically. In response to this, attacks on LLMs are a significant threat to such protections, and many previous approaches have already demonstrated their effectiveness across diverse domains. Existing LLM attacks mostly adopt a generate-and-test strategy to craft malicious input. To improve the comprehension of censoring mechanisms and design a targeted attack, we propose an Explainable-AI solution that comparatively analyzes the behavior of censored and uncensored models to derive unique exploitable alignment patterns. Then, we propose XBreaking, a novel approach that exploits these unique patterns to break the security and alignment constraints of LLMs by targeted noise injection. Our thorough experimental campaign returns important insights about the censoring mechanisms and demonstrates the effectiveness and performance of our approach.
title XBreaking: Understanding how LLMs security alignment can be broken
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2504.21700