State-Dependent Safety Failures in Multi-Turn Language Model Interaction

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Pengcheng, Zhang, Jie, Zhang, Tianwei, Qiu, Han, kejun, Zhang, Zhang, Weiming, Yu, Nenghai, Zhou, Wenbo
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915867829207040
author Li, Pengcheng
Zhang, Jie
Zhang, Tianwei
Qiu, Han
kejun, Zhang
Zhang, Weiming
Yu, Nenghai
Zhou, Wenbo
author_facet Li, Pengcheng
Zhang, Jie
Zhang, Tianwei
Qiu, Han
kejun, Zhang
Zhang, Weiming
Yu, Nenghai
Zhou, Wenbo
contents Safety alignment in large language models is typically evaluated under isolated queries, yet real-world use is inherently multi-turn. Although multi-turn jailbreaks are empirically effective, the structure of conversational safety failure remains insufficiently understood. In this work, we study safety failures from a state-space perspective and show that many multi-turn failures arise from structured contextual state evolution rather than isolated prompt vulnerabilities. We introduce STAR, a state-oriented diagnostic framework that treats dialogue history as a state transition operator and enables controlled analysis of safety behavior along interaction trajectories. Rather than optimizing attack strength, STAR provides a principled probe of how aligned models traverse the safety boundary under autoregressive conditioning. Across multiple frontier language models, we find that systems that appear robust under static evaluation can undergo rapid and reproducible safety collapse under structured multi-turn interaction. Mechanistic analysis reveals monotonic drift away from refusal-related representations and abrupt phase transitions induced by role-conditioned context. Together, these findings motivate viewing language model safety as a dynamic, state-dependent process defined over conversational trajectories.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15684
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle State-Dependent Safety Failures in Multi-Turn Language Model Interaction
Li, Pengcheng
Zhang, Jie
Zhang, Tianwei
Qiu, Han
kejun, Zhang
Zhang, Weiming
Yu, Nenghai
Zhou, Wenbo
Cryptography and Security
Artificial Intelligence
Safety alignment in large language models is typically evaluated under isolated queries, yet real-world use is inherently multi-turn. Although multi-turn jailbreaks are empirically effective, the structure of conversational safety failure remains insufficiently understood. In this work, we study safety failures from a state-space perspective and show that many multi-turn failures arise from structured contextual state evolution rather than isolated prompt vulnerabilities. We introduce STAR, a state-oriented diagnostic framework that treats dialogue history as a state transition operator and enables controlled analysis of safety behavior along interaction trajectories. Rather than optimizing attack strength, STAR provides a principled probe of how aligned models traverse the safety boundary under autoregressive conditioning. Across multiple frontier language models, we find that systems that appear robust under static evaluation can undergo rapid and reproducible safety collapse under structured multi-turn interaction. Mechanistic analysis reveals monotonic drift away from refusal-related representations and abrupt phase transitions induced by role-conditioned context. Together, these findings motivate viewing language model safety as a dynamic, state-dependent process defined over conversational trajectories.
title State-Dependent Safety Failures in Multi-Turn Language Model Interaction
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2603.15684