IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866918383034826752 |
|---|---|
| author | Guo, Chuan Uribe, Juan Felipe Ceron Zhu, Sicheng Choquette-Choo, Christopher A. Lin, Steph Kandpal, Nikhil Nasr, Milad Rai Toyer, Sam Wang, Miles Yu, Yaodong Beutel, Alex Xiao, Kai |
| author_facet | Guo, Chuan Uribe, Juan Felipe Ceron Zhu, Sicheng Choquette-Choo, Christopher A. Lin, Steph Kandpal, Nikhil Nasr, Milad Rai Toyer, Sam Wang, Miles Yu, Yaodong Beutel, Alex Xiao, Kai |
| contents | Instruction hierarchy (IH) defines how LLMs prioritize system, developer, user, and tool instructions under conflict, providing a concrete, trust-ordered policy for resolving instruction conflicts. IH is key to defending against jailbreaks, system prompt extractions, and agentic prompt injections. However, robust IH behavior is difficult to train: IH failures can be confounded with instruction-following failures, conflicts can be nuanced, and models can learn shortcuts such as overrefusing. We introduce IH-Challenge, a reinforcement learning training dataset, to address these difficulties. Fine-tuning GPT-5-Mini on IH-Challenge with online adversarial example generation improves IH robustness by +10.0% on average across 16 in-distribution, out-of-distribution, and human red-teaming benchmarks (84.1% to 94.1%), reduces unsafe behavior from 6.6% to 0.7% while improving helpfulness on general safety evaluations, and saturates an internal static agentic prompt injection evaluation, with minimal capability regression. We release the IH-Challenge dataset (https://huggingface.co/datasets/openai/ih-challenge) to support future research on robust instruction hierarchy. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_10521 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs Guo, Chuan Uribe, Juan Felipe Ceron Zhu, Sicheng Choquette-Choo, Christopher A. Lin, Steph Kandpal, Nikhil Nasr, Milad Rai Toyer, Sam Wang, Miles Yu, Yaodong Beutel, Alex Xiao, Kai Artificial Intelligence Computation and Language Cryptography and Security Machine Learning Instruction hierarchy (IH) defines how LLMs prioritize system, developer, user, and tool instructions under conflict, providing a concrete, trust-ordered policy for resolving instruction conflicts. IH is key to defending against jailbreaks, system prompt extractions, and agentic prompt injections. However, robust IH behavior is difficult to train: IH failures can be confounded with instruction-following failures, conflicts can be nuanced, and models can learn shortcuts such as overrefusing. We introduce IH-Challenge, a reinforcement learning training dataset, to address these difficulties. Fine-tuning GPT-5-Mini on IH-Challenge with online adversarial example generation improves IH robustness by +10.0% on average across 16 in-distribution, out-of-distribution, and human red-teaming benchmarks (84.1% to 94.1%), reduces unsafe behavior from 6.6% to 0.7% while improving helpfulness on general safety evaluations, and saturates an internal static agentic prompt injection evaluation, with minimal capability regression. We release the IH-Challenge dataset (https://huggingface.co/datasets/openai/ih-challenge) to support future research on robust instruction hierarchy. |
| title | IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs |
| topic | Artificial Intelligence Computation and Language Cryptography and Security Machine Learning |
| url | https://arxiv.org/abs/2603.10521 |