IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Guo, Chuan, Uribe, Juan Felipe Ceron, Zhu, Sicheng, Choquette-Choo, Christopher A., Lin, Steph, Kandpal, Nikhil, Nasr, Milad, Rai, Toyer, Sam, Wang, Miles, Yu, Yaodong, Beutel, Alex, Xiao, Kai
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918383034826752
author Guo, Chuan
Uribe, Juan Felipe Ceron
Zhu, Sicheng
Choquette-Choo, Christopher A.
Lin, Steph
Kandpal, Nikhil
Nasr, Milad
Rai
Toyer, Sam
Wang, Miles
Yu, Yaodong
Beutel, Alex
Xiao, Kai
author_facet Guo, Chuan
Uribe, Juan Felipe Ceron
Zhu, Sicheng
Choquette-Choo, Christopher A.
Lin, Steph
Kandpal, Nikhil
Nasr, Milad
Rai
Toyer, Sam
Wang, Miles
Yu, Yaodong
Beutel, Alex
Xiao, Kai
contents Instruction hierarchy (IH) defines how LLMs prioritize system, developer, user, and tool instructions under conflict, providing a concrete, trust-ordered policy for resolving instruction conflicts. IH is key to defending against jailbreaks, system prompt extractions, and agentic prompt injections. However, robust IH behavior is difficult to train: IH failures can be confounded with instruction-following failures, conflicts can be nuanced, and models can learn shortcuts such as overrefusing. We introduce IH-Challenge, a reinforcement learning training dataset, to address these difficulties. Fine-tuning GPT-5-Mini on IH-Challenge with online adversarial example generation improves IH robustness by +10.0% on average across 16 in-distribution, out-of-distribution, and human red-teaming benchmarks (84.1% to 94.1%), reduces unsafe behavior from 6.6% to 0.7% while improving helpfulness on general safety evaluations, and saturates an internal static agentic prompt injection evaluation, with minimal capability regression. We release the IH-Challenge dataset (https://huggingface.co/datasets/openai/ih-challenge) to support future research on robust instruction hierarchy.
format Preprint
id arxiv_https___arxiv_org_abs_2603_10521
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
Guo, Chuan
Uribe, Juan Felipe Ceron
Zhu, Sicheng
Choquette-Choo, Christopher A.
Lin, Steph
Kandpal, Nikhil
Nasr, Milad
Rai
Toyer, Sam
Wang, Miles
Yu, Yaodong
Beutel, Alex
Xiao, Kai
Artificial Intelligence
Computation and Language
Cryptography and Security
Machine Learning
Instruction hierarchy (IH) defines how LLMs prioritize system, developer, user, and tool instructions under conflict, providing a concrete, trust-ordered policy for resolving instruction conflicts. IH is key to defending against jailbreaks, system prompt extractions, and agentic prompt injections. However, robust IH behavior is difficult to train: IH failures can be confounded with instruction-following failures, conflicts can be nuanced, and models can learn shortcuts such as overrefusing. We introduce IH-Challenge, a reinforcement learning training dataset, to address these difficulties. Fine-tuning GPT-5-Mini on IH-Challenge with online adversarial example generation improves IH robustness by +10.0% on average across 16 in-distribution, out-of-distribution, and human red-teaming benchmarks (84.1% to 94.1%), reduces unsafe behavior from 6.6% to 0.7% while improving helpfulness on general safety evaluations, and saturates an internal static agentic prompt injection evaluation, with minimal capability regression. We release the IH-Challenge dataset (https://huggingface.co/datasets/openai/ih-challenge) to support future research on robust instruction hierarchy.
title IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
topic Artificial Intelligence
Computation and Language
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2603.10521