Control Illusion: The Failure of Instruction Hierarchies in Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Geng, Yilin, Li, Haonan, Mu, Honglin, Han, Xudong, Baldwin, Timothy, Abend, Omri, Hovy, Eduard, Frermann, Lea
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908901589385216
author Geng, Yilin
Li, Haonan
Mu, Honglin
Han, Xudong
Baldwin, Timothy
Abend, Omri
Hovy, Eduard
Frermann, Lea
author_facet Geng, Yilin
Li, Haonan
Mu, Honglin
Han, Xudong
Baldwin, Timothy
Abend, Omri
Hovy, Eduard
Frermann, Lea
contents Large language models (LLMs) are increasingly deployed with hierarchical instruction schemes, where certain instructions (e.g., system-level directives) are expected to take precedence over others (e.g., user messages). Yet, we lack a systematic understanding of how effectively these hierarchical control mechanisms work. We introduce a systematic evaluation framework based on constraint prioritization to assess how well LLMs enforce instruction hierarchies. Our experiments across six state-of-the-art LLMs reveal that models struggle with consistent instruction prioritization, even for simple formatting conflicts. We find that the widely-adopted system/user prompt separation fails to establish a reliable instruction hierarchy, and models exhibit strong inherent biases toward certain constraint types regardless of their priority designation. Interestingly, we also find that societal hierarchy framings (e.g., authority, expertise, consensus) show stronger influence on model behavior than system/user roles, suggesting that pretraining-derived social structures function as latent behavioral priors with potentially greater impact than post-training guardrails.
format Preprint
id arxiv_https___arxiv_org_abs_2502_15851
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Control Illusion: The Failure of Instruction Hierarchies in Large Language Models
Geng, Yilin
Li, Haonan
Mu, Honglin
Han, Xudong
Baldwin, Timothy
Abend, Omri
Hovy, Eduard
Frermann, Lea
Computation and Language
Artificial Intelligence
Large language models (LLMs) are increasingly deployed with hierarchical instruction schemes, where certain instructions (e.g., system-level directives) are expected to take precedence over others (e.g., user messages). Yet, we lack a systematic understanding of how effectively these hierarchical control mechanisms work. We introduce a systematic evaluation framework based on constraint prioritization to assess how well LLMs enforce instruction hierarchies. Our experiments across six state-of-the-art LLMs reveal that models struggle with consistent instruction prioritization, even for simple formatting conflicts. We find that the widely-adopted system/user prompt separation fails to establish a reliable instruction hierarchy, and models exhibit strong inherent biases toward certain constraint types regardless of their priority designation. Interestingly, we also find that societal hierarchy framings (e.g., authority, expertise, consensus) show stronger influence on model behavior than system/user roles, suggesting that pretraining-derived social structures function as latent behavioral priors with potentially greater impact than post-training guardrails.
title Control Illusion: The Failure of Instruction Hierarchies in Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2502.15851