Foot-In-The-Door: A Multi-turn Jailbreak for LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Weng, Zixuan, Jin, Xiaolong, Jia, Jinyuan, Zhang, Xiangyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913763076079616
author Weng, Zixuan
Jin, Xiaolong
Jia, Jinyuan
Zhang, Xiangyu
author_facet Weng, Zixuan
Jin, Xiaolong
Jia, Jinyuan
Zhang, Xiangyu
contents Ensuring AI safety is crucial as large language models become increasingly integrated into real-world applications. A key challenge is jailbreak, where adversarial prompts bypass built-in safeguards to elicit harmful disallowed outputs. Inspired by psychological foot-in-the-door principles, we introduce FITD,a novel multi-turn jailbreak method that leverages the phenomenon where minor initial commitments lower resistance to more significant or more unethical transgressions. Our approach progressively escalates the malicious intent of user queries through intermediate bridge prompts and aligns the model's response by itself to induce toxic responses. Extensive experimental results on two jailbreak benchmarks demonstrate that FITD achieves an average attack success rate of 94% across seven widely used models, outperforming existing state-of-the-art methods. Additionally, we provide an in-depth analysis of LLM self-corruption, highlighting vulnerabilities in current alignment strategies and emphasizing the risks inherent in multi-turn interactions. The code is available at https://github.com/Jinxiaolong1129/Foot-in-the-door-Jailbreak.
format Preprint
id arxiv_https___arxiv_org_abs_2502_19820
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Foot-In-The-Door: A Multi-turn Jailbreak for LLMs
Weng, Zixuan
Jin, Xiaolong
Jia, Jinyuan
Zhang, Xiangyu
Computation and Language
Artificial Intelligence
Ensuring AI safety is crucial as large language models become increasingly integrated into real-world applications. A key challenge is jailbreak, where adversarial prompts bypass built-in safeguards to elicit harmful disallowed outputs. Inspired by psychological foot-in-the-door principles, we introduce FITD,a novel multi-turn jailbreak method that leverages the phenomenon where minor initial commitments lower resistance to more significant or more unethical transgressions. Our approach progressively escalates the malicious intent of user queries through intermediate bridge prompts and aligns the model's response by itself to induce toxic responses. Extensive experimental results on two jailbreak benchmarks demonstrate that FITD achieves an average attack success rate of 94% across seven widely used models, outperforming existing state-of-the-art methods. Additionally, we provide an in-depth analysis of LLM self-corruption, highlighting vulnerabilities in current alignment strategies and emphasizing the risks inherent in multi-turn interactions. The code is available at https://github.com/Jinxiaolong1129/Foot-in-the-door-Jailbreak.
title Foot-In-The-Door: A Multi-turn Jailbreak for LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2502.19820