Internal Safety Collapse in Frontier Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yutao, Liu, Xiao, Gao, Yifeng, Zheng, Xiang, Huang, Hanxun, Li, Yige, Wang, Cong, Li, Bo, Ma, Xingjun, Jiang, Yu-Gang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911542764634112
author Wu, Yutao
Liu, Xiao
Gao, Yifeng
Zheng, Xiang
Huang, Hanxun
Li, Yige
Wang, Cong
Li, Bo
Ma, Xingjun
Jiang, Yu-Gang
author_facet Wu, Yutao
Liu, Xiao
Gao, Yifeng
Zheng, Xiang
Huang, Hanxun
Li, Yige
Wang, Cong
Li, Bo
Ma, Xingjun
Jiang, Yu-Gang
contents This work identifies a critical failure mode in frontier large language models (LLMs), which we term Internal Safety Collapse (ISC): under certain task conditions, models enter a state in which they continuously generate harmful content while executing otherwise benign tasks. We introduce TVD (Task, Validator, Data), a framework that triggers ISC through domain tasks where generating harmful content is the only valid completion, and construct ISC-Bench containing 53 scenarios across 8 professional disciplines. Evaluated on JailbreakBench, three representative scenarios yield worst-case safety failure rates averaging 95.3% across four frontier LLMs (including GPT-5.2 and Claude Sonnet 4.5), substantially exceeding standard jailbreak attacks. Frontier models are more vulnerable than earlier LLMs: the very capabilities that enable complex task execution become liabilities when tasks intrinsically involve harmful content. This reveals a growing attack surface: almost every professional domain uses tools that process sensitive data, and each new dual-use tool automatically expands this vulnerability--even without any deliberate attack. Despite substantial alignment efforts, frontier LLMs retain inherently unsafe internal capabilities: alignment reshapes observable outputs but does not eliminate the underlying risk profile. These findings underscore the need for caution when deploying LLMs in high-stakes settings. Source code: https://github.com/wuyoscar/ISC-Bench
format Preprint
id arxiv_https___arxiv_org_abs_2603_23509
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Internal Safety Collapse in Frontier Large Language Models
Wu, Yutao
Liu, Xiao
Gao, Yifeng
Zheng, Xiang
Huang, Hanxun
Li, Yige
Wang, Cong
Li, Bo
Ma, Xingjun
Jiang, Yu-Gang
Computation and Language
Artificial Intelligence
Cryptography and Security
This work identifies a critical failure mode in frontier large language models (LLMs), which we term Internal Safety Collapse (ISC): under certain task conditions, models enter a state in which they continuously generate harmful content while executing otherwise benign tasks. We introduce TVD (Task, Validator, Data), a framework that triggers ISC through domain tasks where generating harmful content is the only valid completion, and construct ISC-Bench containing 53 scenarios across 8 professional disciplines. Evaluated on JailbreakBench, three representative scenarios yield worst-case safety failure rates averaging 95.3% across four frontier LLMs (including GPT-5.2 and Claude Sonnet 4.5), substantially exceeding standard jailbreak attacks. Frontier models are more vulnerable than earlier LLMs: the very capabilities that enable complex task execution become liabilities when tasks intrinsically involve harmful content. This reveals a growing attack surface: almost every professional domain uses tools that process sensitive data, and each new dual-use tool automatically expands this vulnerability--even without any deliberate attack. Despite substantial alignment efforts, frontier LLMs retain inherently unsafe internal capabilities: alignment reshapes observable outputs but does not eliminate the underlying risk profile. These findings underscore the need for caution when deploying LLMs in high-stakes settings. Source code: https://github.com/wuyoscar/ISC-Bench
title Internal Safety Collapse in Frontier Large Language Models
topic Computation and Language
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2603.23509