One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Xinjie, Wei, Rongzhe, Niu, Peizhi, Wang, Haoyu, Wu, Ruihan, Chien, Eli, Li, Bo, Chen, Pin-Yu, Li, Pan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911675142111232
author Shen, Xinjie
Wei, Rongzhe
Niu, Peizhi
Wang, Haoyu
Wu, Ruihan
Chien, Eli
Li, Bo
Chen, Pin-Yu
Li, Pan
author_facet Shen, Xinjie
Wei, Rongzhe
Niu, Peizhi
Wang, Haoyu
Wu, Ruihan
Chien, Eli
Li, Bo
Chen, Pin-Yu
Li, Pan
contents Hidden malicious intent in multi-turn dialogue poses a growing threat to deployed large language models (LLMs). Rather than exposing a harmful objective in a single prompt, increasingly capable attackers can distribute their intent across multiple benign-looking turns. Recent studies show that even modern commercial models with advanced guardrails remain vulnerable to such attacks despite advances in safety alignment and external guardrails. In this work, we address this challenge by detecting the earliest turn at which delivering the candidate response would make the accumulated interaction sufficient to enable harmful action. This objective requires precise turn-level intervention that identifies the harm-enabling closure point while avoiding premature refusal of benign exploratory conversations. To further support training and evaluation, we construct the Multi-Turn Intent Dataset (MTID), which contains branching attack rollouts, matched benign hard negatives, and annotations of the earliest harm-enabling turns. We show that MTID helps enable a turn-level monitor TurnGate, which substantially outperforms existing baselines in harmful-intent detection while maintaining low over-refusal rates. TurnGate further generalizes across domains, attacker pipelines, and target models. Our code is available at https://github.com/Graph-COM/TurnGate.
format Preprint
id arxiv_https___arxiv_org_abs_2605_05630
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
Shen, Xinjie
Wei, Rongzhe
Niu, Peizhi
Wang, Haoyu
Wu, Ruihan
Chien, Eli
Li, Bo
Chen, Pin-Yu
Li, Pan
Computation and Language
Artificial Intelligence
Cryptography and Security
Hidden malicious intent in multi-turn dialogue poses a growing threat to deployed large language models (LLMs). Rather than exposing a harmful objective in a single prompt, increasingly capable attackers can distribute their intent across multiple benign-looking turns. Recent studies show that even modern commercial models with advanced guardrails remain vulnerable to such attacks despite advances in safety alignment and external guardrails. In this work, we address this challenge by detecting the earliest turn at which delivering the candidate response would make the accumulated interaction sufficient to enable harmful action. This objective requires precise turn-level intervention that identifies the harm-enabling closure point while avoiding premature refusal of benign exploratory conversations. To further support training and evaluation, we construct the Multi-Turn Intent Dataset (MTID), which contains branching attack rollouts, matched benign hard negatives, and annotations of the earliest harm-enabling turns. We show that MTID helps enable a turn-level monitor TurnGate, which substantially outperforms existing baselines in harmful-intent detection while maintaining low over-refusal rates. TurnGate further generalizes across domains, attacker pipelines, and target models. Our code is available at https://github.com/Graph-COM/TurnGate.
title One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
topic Computation and Language
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2605.05630