AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Jiacheng, Jiang, Tanqiu, Wang, Yuhui, Zhu, Rongyi, Ma, Fenglong, Wang, Ting
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914477830569984
author Liang, Jiacheng
Jiang, Tanqiu
Wang, Yuhui
Zhu, Rongyi
Ma, Fenglong
Wang, Ting
author_facet Liang, Jiacheng
Jiang, Tanqiu
Wang, Yuhui
Zhu, Rongyi
Ma, Fenglong
Wang, Ting
contents This paper presents AutoRAN, the first framework to automate the hijacking of internal safety reasoning in large reasoning models (LRMs). At its core, AutoRAN pioneers an execution simulation paradigm that leverages a weaker but less-aligned model to simulate execution reasoning for initial hijacking attempts and iteratively refine attacks by exploiting reasoning patterns leaked through the target LRM's refusals. This approach steers the target model to bypass its own safety guardrails and elaborate on harmful instructions. We evaluate AutoRAN against state-of-the-art LRMs, including GPT-o3/o4-mini and Gemini-2.5-Flash, across multiple benchmarks (AdvBench, HarmBench, and StrongReject). Results show that AutoRAN achieves approaching 100% success rate within one or few turns, effectively neutralizing reasoning-based defenses even when evaluated by robustly aligned external models. This work reveals that the transparency of the reasoning process itself creates a critical and exploitable attack surface, highlighting the urgent need for new defenses that protect models' reasoning traces rather than merely their final outputs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_10846
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models
Liang, Jiacheng
Jiang, Tanqiu
Wang, Yuhui
Zhu, Rongyi
Ma, Fenglong
Wang, Ting
Machine Learning
Cryptography and Security
This paper presents AutoRAN, the first framework to automate the hijacking of internal safety reasoning in large reasoning models (LRMs). At its core, AutoRAN pioneers an execution simulation paradigm that leverages a weaker but less-aligned model to simulate execution reasoning for initial hijacking attempts and iteratively refine attacks by exploiting reasoning patterns leaked through the target LRM's refusals. This approach steers the target model to bypass its own safety guardrails and elaborate on harmful instructions. We evaluate AutoRAN against state-of-the-art LRMs, including GPT-o3/o4-mini and Gemini-2.5-Flash, across multiple benchmarks (AdvBench, HarmBench, and StrongReject). Results show that AutoRAN achieves approaching 100% success rate within one or few turns, effectively neutralizing reasoning-based defenses even when evaluated by robustly aligned external models. This work reveals that the transparency of the reasoning process itself creates a critical and exploitable attack surface, highlighting the urgent need for new defenses that protect models' reasoning traces rather than merely their final outputs.
title AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2505.10846