Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yueh-Han, Chen, Joshi, Nitish, Chen, Yulin, Andriushchenko, Maksym, Angell, Rico, He, He
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915342900527104
author Yueh-Han, Chen
Joshi, Nitish
Chen, Yulin
Andriushchenko, Maksym
Angell, Rico
He, He
author_facet Yueh-Han, Chen
Joshi, Nitish
Chen, Yulin
Andriushchenko, Maksym
Angell, Rico
He, He
contents Current LLM safety defenses fail under decomposition attacks, where a malicious goal is decomposed into benign subtasks that circumvent refusals. The challenge lies in the existing shallow safety alignment techniques: they only detect harm in the immediate prompt and do not reason about long-range intent, leaving them blind to malicious intent that emerges over a sequence of seemingly benign instructions. We therefore propose adding an external monitor that observes the conversation at a higher granularity. To facilitate our study of monitoring decomposition attacks, we curate the largest and most diverse dataset to date, including question-answering, text-to-image, and agentic tasks. We verify our datasets by testing them on frontier LLMs and show an 87% attack success rate on average on GPT-4o. This confirms that decomposition attack is broadly effective. Additionally, we find that random tasks can be injected into the decomposed subtasks to further obfuscate malicious intents. To defend in real time, we propose a lightweight sequential monitoring framework that cumulatively evaluates each subtask. We show that a carefully prompt engineered lightweight monitor achieves a 93% defense success rate, beating reasoning models like o3 mini as a monitor. Moreover, it remains robust against random task injection and cuts cost by 90% and latency by 50%. Our findings suggest that lightweight sequential monitors are highly effective in mitigating decomposition attacks and are viable in deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10949
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors
Yueh-Han, Chen
Joshi, Nitish
Chen, Yulin
Andriushchenko, Maksym
Angell, Rico
He, He
Cryptography and Security
Artificial Intelligence
Current LLM safety defenses fail under decomposition attacks, where a malicious goal is decomposed into benign subtasks that circumvent refusals. The challenge lies in the existing shallow safety alignment techniques: they only detect harm in the immediate prompt and do not reason about long-range intent, leaving them blind to malicious intent that emerges over a sequence of seemingly benign instructions. We therefore propose adding an external monitor that observes the conversation at a higher granularity. To facilitate our study of monitoring decomposition attacks, we curate the largest and most diverse dataset to date, including question-answering, text-to-image, and agentic tasks. We verify our datasets by testing them on frontier LLMs and show an 87% attack success rate on average on GPT-4o. This confirms that decomposition attack is broadly effective. Additionally, we find that random tasks can be injected into the decomposed subtasks to further obfuscate malicious intents. To defend in real time, we propose a lightweight sequential monitoring framework that cumulatively evaluates each subtask. We show that a carefully prompt engineered lightweight monitor achieves a 93% defense success rate, beating reasoning models like o3 mini as a monitor. Moreover, it remains robust against random task injection and cuts cost by 90% and latency by 50%. Our findings suggest that lightweight sequential monitors are highly effective in mitigating decomposition attacks and are viable in deployment.
title Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2506.10949