Path Drift in Large Reasoning Models:How First-Person Commitments Override Safety

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Huang, Yuyi, Zhan, Runzhe, Chao, Lidia S., Tao, Ailin, Wong, Derek F.
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908587578621952
author Huang, Yuyi
Zhan, Runzhe
Chao, Lidia S.
Tao, Ailin
Wong, Derek F.
author_facet Huang, Yuyi
Zhan, Runzhe
Chao, Lidia S.
Tao, Ailin
Wong, Derek F.
contents As large language models (LLMs) are increasingly deployed for complex reasoning tasks, Long Chain-of-Thought (Long-CoT) prompting has emerged as a key paradigm for structured inference. Despite early-stage safeguards enabled by alignment techniques such as RLHF, we identify a previously underexplored vulnerability: reasoning trajectories in Long-CoT models can drift from aligned paths, resulting in content that violates safety constraints. We term this phenomenon Path Drift. Through empirical analysis, we uncover three behavioral triggers of Path Drift: (1) first-person commitments that induce goal-driven reasoning that delays refusal signals; (2) ethical evaporation, where surface-level disclaimers bypass alignment checkpoints; (3) condition chain escalation, where layered cues progressively steer models toward unsafe completions. Building on these insights, we introduce a three-stage Path Drift Induction Framework comprising cognitive load amplification, self-role priming, and condition chain hijacking. Each stage independently reduces refusal rates, while their combination further compounds the effect. To mitigate these risks, we propose a path-level defense strategy incorporating role attribution correction and metacognitive reflection (reflective safety cues). Our findings highlight the need for trajectory-level alignment oversight in long-form reasoning beyond token-level alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10013
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Path Drift in Large Reasoning Models:How First-Person Commitments Override Safety
Huang, Yuyi
Zhan, Runzhe
Chao, Lidia S.
Tao, Ailin
Wong, Derek F.
Computation and Language
As large language models (LLMs) are increasingly deployed for complex reasoning tasks, Long Chain-of-Thought (Long-CoT) prompting has emerged as a key paradigm for structured inference. Despite early-stage safeguards enabled by alignment techniques such as RLHF, we identify a previously underexplored vulnerability: reasoning trajectories in Long-CoT models can drift from aligned paths, resulting in content that violates safety constraints. We term this phenomenon Path Drift. Through empirical analysis, we uncover three behavioral triggers of Path Drift: (1) first-person commitments that induce goal-driven reasoning that delays refusal signals; (2) ethical evaporation, where surface-level disclaimers bypass alignment checkpoints; (3) condition chain escalation, where layered cues progressively steer models toward unsafe completions. Building on these insights, we introduce a three-stage Path Drift Induction Framework comprising cognitive load amplification, self-role priming, and condition chain hijacking. Each stage independently reduces refusal rates, while their combination further compounds the effect. To mitigate these risks, we propose a path-level defense strategy incorporating role attribution correction and metacognitive reflection (reflective safety cues). Our findings highlight the need for trajectory-level alignment oversight in long-form reasoning beyond token-level alignment.
title Path Drift in Large Reasoning Models:How First-Person Commitments Override Safety
topic Computation and Language
url https://arxiv.org/abs/2510.10013