Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xunguang, Zhou, Yuguang, Wang, Qingyue, Li, Zongjie, Huang, Ruixuan, Ji, Zhenlan, Ma, Pingchuan, Wang, Shuai
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917462838083584
author Wang, Xunguang
Zhou, Yuguang
Wang, Qingyue
Li, Zongjie
Huang, Ruixuan
Ji, Zhenlan
Ma, Pingchuan
Wang, Shuai
author_facet Wang, Xunguang
Zhou, Yuguang
Wang, Qingyue
Li, Zongjie
Huang, Ruixuan
Ji, Zhenlan
Ma, Pingchuan
Wang, Shuai
contents Large language models increasingly rely on explicit chain-of-thought reasoning to solve complex tasks, yet the safety of the reasoning process itself remains largely unaddressed. Existing work focuses predominantly on content safety (i.e., detecting harmful, biased, or factually incorrect outputs), while treating the underlying reasoning chain as an opaque intermediate artifact. We argue that reasoning safety constitutes a fundamental security dimension orthogonal to content safety: the requirement that a model's reasoning trajectory be logically consistent, computationally efficient, and resistant to adversarial manipulation. In this paper, we formalize reasoning safety and introduce a systematic taxonomy of nine unsafe reasoning behaviors. We then conduct a large-scale prevalence study, annotating over 4,000 reasoning chains across benign benchmarks and four state-of-the-art reasoning attacks, empirically demonstrating that all nine error types occur in practice with mechanistically interpretable signatures. To mitigate these threats, we propose the Reasoning Safety Monitor: an external, zero-shot verification framework that runs in parallel with the target LLM. It inspects each reasoning step in real time via a taxonomy-embedded prompt and dispatches an interrupt signal upon detecting unsafe behavior. Extensive evaluations show our monitor achieves up to 87.11% step-level localization accuracy, outperforming hallucination detectors and the best process reward model baselines by a substantial margin. Crucially, the monitor maintains a low false positive rate on correct reasoning paths, operates with negligible latency overhead, and exhibits robust resilience against adaptive adversarial evasion. These findings establish reasoning safety monitoring as a highly feasible and essential component for the secure deployment of large reasoning models.
format Preprint
id arxiv_https___arxiv_org_abs_2603_25412
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models
Wang, Xunguang
Zhou, Yuguang
Wang, Qingyue
Li, Zongjie
Huang, Ruixuan
Ji, Zhenlan
Ma, Pingchuan
Wang, Shuai
Artificial Intelligence
Cryptography and Security
Large language models increasingly rely on explicit chain-of-thought reasoning to solve complex tasks, yet the safety of the reasoning process itself remains largely unaddressed. Existing work focuses predominantly on content safety (i.e., detecting harmful, biased, or factually incorrect outputs), while treating the underlying reasoning chain as an opaque intermediate artifact. We argue that reasoning safety constitutes a fundamental security dimension orthogonal to content safety: the requirement that a model's reasoning trajectory be logically consistent, computationally efficient, and resistant to adversarial manipulation. In this paper, we formalize reasoning safety and introduce a systematic taxonomy of nine unsafe reasoning behaviors. We then conduct a large-scale prevalence study, annotating over 4,000 reasoning chains across benign benchmarks and four state-of-the-art reasoning attacks, empirically demonstrating that all nine error types occur in practice with mechanistically interpretable signatures. To mitigate these threats, we propose the Reasoning Safety Monitor: an external, zero-shot verification framework that runs in parallel with the target LLM. It inspects each reasoning step in real time via a taxonomy-embedded prompt and dispatches an interrupt signal upon detecting unsafe behavior. Extensive evaluations show our monitor achieves up to 87.11% step-level localization accuracy, outperforming hallucination detectors and the best process reward model baselines by a substantial margin. Crucially, the monitor maintains a low false positive rate on correct reasoning paths, operates with negligible latency overhead, and exhibits robust resilience against adaptive adversarial evasion. These findings establish reasoning safety monitoring as a highly feasible and essential component for the secure deployment of large reasoning models.
title Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models
topic Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2603.25412