Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Huilin, Zhao, Jian, Zhong, Yilu, Liang, Zhen, Chen, Xiuyuan, Yuan, Yuchen, Zhang, Tianle, Zhang, Chi, Zhang, Lan, Li, Xuelong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914586775519232
author Zhou, Huilin
Zhao, Jian
Zhong, Yilu
Liang, Zhen
Chen, Xiuyuan
Yuan, Yuchen
Zhang, Tianle
Zhang, Chi
Zhang, Lan
Li, Xuelong
author_facet Zhou, Huilin
Zhao, Jian
Zhong, Yilu
Liang, Zhen
Chen, Xiuyuan
Yuan, Yuchen
Zhang, Tianle
Zhang, Chi
Zhang, Lan
Li, Xuelong
contents Red teaming is critical for uncovering vulnerabilities in Large Language Models (LLMs). While automated methods have improved scalability, existing approaches often rely on static heuristics or stochastic search, rendering them brittle against advanced safety alignment. To address this, we introduce Metis, a framework that reformulates jailbreaking as inference-time policy optimization within an adversarial Partially Observable Markov Decision Process (POMDP). Metis employs a self-evolving metacognitive loop to perform causal diagnosis of a target's defense logic and leverages structured feedback as a semantic gradient to refine its policy, offering enhanced interpretability through transparent reasoning traces. Extensive evaluations across 10 diverse models demonstrate that Metis achieves the strongest average Attack Success Rate (ASR) among compared methods at 89.2%, maintaining high efficacy on resilient frontier models (e.g., 76.0% on O1 and 78.0% on GPT-5-chat) where traditional baselines exhibit substantial performance degradation. By replacing redundant exploration with directed optimization, Metis reduces token costs by an average of 8.2x and up to 11.4x. Our analysis reveals that current defenses remain vulnerable to internally-steered, closed-loop reasoning trajectories under the tested settings, highlighting a critical need for next-generation defenses capable of reasoning about safety dynamically during inference.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10067
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization
Zhou, Huilin
Zhao, Jian
Zhong, Yilu
Liang, Zhen
Chen, Xiuyuan
Yuan, Yuchen
Zhang, Tianle
Zhang, Chi
Zhang, Lan
Li, Xuelong
Machine Learning
Artificial Intelligence
Red teaming is critical for uncovering vulnerabilities in Large Language Models (LLMs). While automated methods have improved scalability, existing approaches often rely on static heuristics or stochastic search, rendering them brittle against advanced safety alignment. To address this, we introduce Metis, a framework that reformulates jailbreaking as inference-time policy optimization within an adversarial Partially Observable Markov Decision Process (POMDP). Metis employs a self-evolving metacognitive loop to perform causal diagnosis of a target's defense logic and leverages structured feedback as a semantic gradient to refine its policy, offering enhanced interpretability through transparent reasoning traces. Extensive evaluations across 10 diverse models demonstrate that Metis achieves the strongest average Attack Success Rate (ASR) among compared methods at 89.2%, maintaining high efficacy on resilient frontier models (e.g., 76.0% on O1 and 78.0% on GPT-5-chat) where traditional baselines exhibit substantial performance degradation. By replacing redundant exploration with directed optimization, Metis reduces token costs by an average of 8.2x and up to 11.4x. Our analysis reveals that current defenses remain vulnerable to internally-steered, closed-loop reasoning trajectories under the tested settings, highlighting a critical need for next-generation defenses capable of reasoning about safety dynamically during inference.
title Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.10067