Exploiting Latent Space Discontinuities for Building Universal LLM Jailbreaks and Data Extraction Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Paim, Kayua Oleques, Mansilha, Rodrigo Brandao, Kreutz, Diego, Franco, Muriel Figueredo, Cordeiro, Weverton
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909880983486464
author Paim, Kayua Oleques
Mansilha, Rodrigo Brandao
Kreutz, Diego
Franco, Muriel Figueredo
Cordeiro, Weverton
author_facet Paim, Kayua Oleques
Mansilha, Rodrigo Brandao
Kreutz, Diego
Franco, Muriel Figueredo
Cordeiro, Weverton
contents The rapid proliferation of Large Language Models (LLMs) has raised significant concerns about their security against adversarial attacks. In this work, we propose a novel approach to crafting universal jailbreaks and data extraction attacks by exploiting latent space discontinuities, an architectural vulnerability related to the sparsity of training data. Unlike previous methods, our technique generalizes across various models and interfaces, proving highly effective in seven state-of-the-art LLMs and one image generation model. Initial results indicate that when these discontinuities are exploited, they can consistently and profoundly compromise model behavior, even in the presence of layered defenses. The findings suggest that this strategy has substantial potential as a systemic attack vector.
format Preprint
id arxiv_https___arxiv_org_abs_2511_00346
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploiting Latent Space Discontinuities for Building Universal LLM Jailbreaks and Data Extraction Attacks
Paim, Kayua Oleques
Mansilha, Rodrigo Brandao
Kreutz, Diego
Franco, Muriel Figueredo
Cordeiro, Weverton
Cryptography and Security
Artificial Intelligence
Machine Learning
68T01
I.2
The rapid proliferation of Large Language Models (LLMs) has raised significant concerns about their security against adversarial attacks. In this work, we propose a novel approach to crafting universal jailbreaks and data extraction attacks by exploiting latent space discontinuities, an architectural vulnerability related to the sparsity of training data. Unlike previous methods, our technique generalizes across various models and interfaces, proving highly effective in seven state-of-the-art LLMs and one image generation model. Initial results indicate that when these discontinuities are exploited, they can consistently and profoundly compromise model behavior, even in the presence of layered defenses. The findings suggest that this strategy has substantial potential as a systemic attack vector.
title Exploiting Latent Space Discontinuities for Building Universal LLM Jailbreaks and Data Extraction Attacks
topic Cryptography and Security
Artificial Intelligence
Machine Learning
68T01
I.2
url https://arxiv.org/abs/2511.00346