Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Hee-Seon, Kim, Minbeom, Lee, Wonjun, Kim, Kihyun, Kim, Changick
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910970983481344
author Kim, Hee-Seon
Kim, Minbeom
Lee, Wonjun
Kim, Kihyun
Kim, Changick
author_facet Kim, Hee-Seon
Kim, Minbeom
Lee, Wonjun
Kim, Kihyun
Kim, Changick
contents Optimization-based jailbreaks typically adopt the Toxic-Continuation setting in large vision-language models (LVLMs), following the standard next-token prediction objective. In this setting, an adversarial image is optimized to make the model predict the next token of a toxic prompt. However, we find that the Toxic-Continuation paradigm is effective at continuing already-toxic inputs, but struggles to induce safety misalignment when explicit toxic signals are absent. We propose a new paradigm: Benign-to-Toxic (B2T) jailbreak. Unlike prior work, we optimize adversarial images to induce toxic outputs from benign conditioning. Since benign conditioning contains no safety violations, the image alone must break the model's safety mechanisms. Our method outperforms prior approaches, transfers in black-box settings, and complements text-based jailbreaks. These results reveal an underexplored vulnerability in multimodal alignment and introduce a fundamentally new direction for jailbreak approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21556
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts
Kim, Hee-Seon
Kim, Minbeom
Lee, Wonjun
Kim, Kihyun
Kim, Changick
Computer Vision and Pattern Recognition
Artificial Intelligence
Optimization-based jailbreaks typically adopt the Toxic-Continuation setting in large vision-language models (LVLMs), following the standard next-token prediction objective. In this setting, an adversarial image is optimized to make the model predict the next token of a toxic prompt. However, we find that the Toxic-Continuation paradigm is effective at continuing already-toxic inputs, but struggles to induce safety misalignment when explicit toxic signals are absent. We propose a new paradigm: Benign-to-Toxic (B2T) jailbreak. Unlike prior work, we optimize adversarial images to induce toxic outputs from benign conditioning. Since benign conditioning contains no safety violations, the image alone must break the model's safety mechanisms. Our method outperforms prior approaches, transfers in black-box settings, and complements text-based jailbreaks. These results reveal an underexplored vulnerability in multimodal alignment and introduce a fundamentally new direction for jailbreak approaches.
title Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.21556