Many-Tier Instruction Hierarchy in LLM Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jingyu, Li, Tianjian, Jurayj, William, Zhan, Hongyuan, Van Durme, Benjamin, Khashabi, Daniel
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915937007960064
author Zhang, Jingyu
Li, Tianjian
Jurayj, William
Zhan, Hongyuan
Van Durme, Benjamin
Khashabi, Daniel
author_facet Zhang, Jingyu
Li, Tianjian
Jurayj, William
Zhan, Hongyuan
Van Durme, Benjamin
Khashabi, Daniel
contents Large language model agents receive instructions from many sources-system messages, user prompts, tool outputs, other agents, and more-each carrying different levels of trust and authority. When these instructions conflict, agents must reliably follow the highest-privilege instruction to remain safe and effective. The dominant paradigm, instruction hierarchy (IH), assumes a fixed, small set of privilege levels (typically fewer than five) defined by rigid role labels (e.g., system > user). This is inadequate for real-world agentic settings, where conflicts can arise across far more sources and contexts. In this work, we propose Many-Tier Instruction Hierarchy (ManyIH), a paradigm for resolving instruction conflicts among instructions with arbitrarily many privilege levels. We introduce ManyIH-Bench, the first benchmark for ManyIH. ManyIH-Bench requires models to navigate up to 12 levels of conflicting instructions with varying privileges, comprising 853 agentic tasks (427 coding and 426 instruction-following). ManyIH-Bench composes constraints developed by LLMs and verified by humans to create realistic and difficult test cases spanning 46 real-world agents. Our experiments show that even the current frontier models perform poorly (~40% accuracy) when instruction conflict scales. This work underscores the urgent need for methods that explicitly target fine-grained, scalable instruction conflict resolution in agentic settings.
format Preprint
id arxiv_https___arxiv_org_abs_2604_09443
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Many-Tier Instruction Hierarchy in LLM Agents
Zhang, Jingyu
Li, Tianjian
Jurayj, William
Zhan, Hongyuan
Van Durme, Benjamin
Khashabi, Daniel
Computation and Language
Artificial Intelligence
Large language model agents receive instructions from many sources-system messages, user prompts, tool outputs, other agents, and more-each carrying different levels of trust and authority. When these instructions conflict, agents must reliably follow the highest-privilege instruction to remain safe and effective. The dominant paradigm, instruction hierarchy (IH), assumes a fixed, small set of privilege levels (typically fewer than five) defined by rigid role labels (e.g., system > user). This is inadequate for real-world agentic settings, where conflicts can arise across far more sources and contexts. In this work, we propose Many-Tier Instruction Hierarchy (ManyIH), a paradigm for resolving instruction conflicts among instructions with arbitrarily many privilege levels. We introduce ManyIH-Bench, the first benchmark for ManyIH. ManyIH-Bench requires models to navigate up to 12 levels of conflicting instructions with varying privileges, comprising 853 agentic tasks (427 coding and 426 instruction-following). ManyIH-Bench composes constraints developed by LLMs and verified by humans to create realistic and difficult test cases spanning 46 real-world agents. Our experiments show that even the current frontier models perform poorly (~40% accuracy) when instruction conflict scales. This work underscores the urgent need for methods that explicitly target fine-grained, scalable instruction conflict resolution in agentic settings.
title Many-Tier Instruction Hierarchy in LLM Agents
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.09443