IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Chuan, Uribe, Juan Felipe Ceron, Zhu, Sicheng, Choquette-Choo, Christopher A., Lin, Steph, Kandpal, Nikhil, Nasr, Milad, Rai, Toyer, Sam, Wang, Miles, Yu, Yaodong, Beutel, Alex, Xiao, Kai |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
by: Wallace, Eric, et al.
Published: (2024)
by: Wallace, Eric, et al.
Published: (2024)
Auditing Private Prediction
by: Chadha, Karan, et al.
Published: (2024)
by: Chadha, Karan, et al.
Published: (2024)
Privacy Auditing of Large Language Models
by: Panda, Ashwinee, et al.
Published: (2025)
by: Panda, Ashwinee, et al.
Published: (2025)
User Inference Attacks on Large Language Models
by: Kandpal, Nikhil, et al.
Published: (2023)
by: Kandpal, Nikhil, et al.
Published: (2023)
LLMs unlock new paths to monetizing exploits
by: Carlini, Nicholas, et al.
Published: (2025)
by: Carlini, Nicholas, et al.
Published: (2025)
Position: The Most Expensive Part of an LLM should be its Training Data
by: Kandpal, Nikhil, et al.
Published: (2025)
by: Kandpal, Nikhil, et al.
Published: (2025)
Privacy Side Channels in Machine Learning Systems
by: Debenedetti, Edoardo, et al.
Published: (2023)
by: Debenedetti, Edoardo, et al.
Published: (2023)
Exploring and Addressing Reward Confusion in Offline Preference Learning
by: Chen, Xin, et al.
Published: (2024)
by: Chen, Xin, et al.
Published: (2024)
AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution
by: Liu, Fengyuan, et al.
Published: (2024)
by: Liu, Fengyuan, et al.
Published: (2024)
Measuring memorization in language models via probabilistic extraction
by: Hayes, Jamie, et al.
Published: (2024)
by: Hayes, Jamie, et al.
Published: (2024)
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
by: Chaudhari, Harsh, et al.
Published: (2024)
by: Chaudhari, Harsh, et al.
Published: (2024)
Teach LLMs to Phish: Stealing Private Information from Language Models
by: Panda, Ashwinee, et al.
Published: (2024)
by: Panda, Ashwinee, et al.
Published: (2024)
Remote Timing Attacks on Efficient Language Model Inference
by: Carlini, Nicholas, et al.
Published: (2024)
by: Carlini, Nicholas, et al.
Published: (2024)
Avoiding Generative Model Writer's Block With Embedding Nudging
by: Zand, Ali, et al.
Published: (2024)
by: Zand, Ali, et al.
Published: (2024)
The Last Iterate Advantage: Empirical Auditing and Principled Heuristic Analysis of Differentially Private SGD
by: Steinke, Thomas, et al.
Published: (2024)
by: Steinke, Thomas, et al.
Published: (2024)
Optimal Rates for $O(1)$-Smooth DP-SCO with a Single Epoch and Large Batches
by: Choquette-Choo, Christopher A., et al.
Published: (2024)
by: Choquette-Choo, Christopher A., et al.
Published: (2024)
Efficient and Privacy-Preserving Soft Prompt Transfer for LLMs
by: Wang, Xun, et al.
Published: (2025)
by: Wang, Xun, et al.
Published: (2025)
Are aligned neural networks adversarially aligned?
by: Carlini, Nicholas, et al.
Published: (2023)
by: Carlini, Nicholas, et al.
Published: (2023)
Challenges and Opportunities for Proton Batteries: From Electrodes, Electrolytes to Full‐Cell Applications
by: Sicheng Wu, et al.
Published: (2024)
by: Sicheng Wu, et al.
Published: (2024)
Privacy Amplification for Matrix Mechanisms
by: Choquette-Choo, Christopher A., et al.
Published: (2023)
by: Choquette-Choo, Christopher A., et al.
Published: (2023)
Estimating Worst-Case Frontier Risks of Open-Weight LLMs
by: Wallace, Eric, et al.
Published: (2025)
by: Wallace, Eric, et al.
Published: (2025)
Poisoning Web-Scale Training Datasets is Practical
by: Carlini, Nicholas, et al.
Published: (2023)
by: Carlini, Nicholas, et al.
Published: (2023)
Enhancing Training Data Attribution with Representational Optimization
by: Sun, Weiwei, et al.
Published: (2025)
by: Sun, Weiwei, et al.
Published: (2025)
Trading Inference-Time Compute for Adversarial Robustness
by: Zaremba, Wojciech, et al.
Published: (2025)
by: Zaremba, Wojciech, et al.
Published: (2025)
A goodness-of-fit test for testing exponentiality based on normalized dynamic survival extropy
by: Kandpal, Gaurav, et al.
Published: (2024)
by: Kandpal, Gaurav, et al.
Published: (2024)
Characterization based Goodness-of-Fit for Generalized Pareto Distribution: A Blend of Stein's Identity and Dynamic Survival Extropy
by: Kandpal, Gaurav, et al.
Published: (2025)
by: Kandpal, Gaurav, et al.
Published: (2025)
A Survey of Body and Face Motion: Datasets, Performance Evaluation Metrics and Generative Techniques
by: Sookha, Lownish Rai, et al.
Published: (2025)
by: Sookha, Lownish Rai, et al.
Published: (2025)
Efficient Model Development through Fine-tuning Transfer
by: Lin, Pin-Jie, et al.
Published: (2025)
by: Lin, Pin-Jie, et al.
Published: (2025)
The Relationship Between Discomfort Intolerance And the Fear Of Self‐Injection And Testing In Patients With Diabetes Using Insulin: A Cross‐Sectional Study
by: Nilhan Töyer Şahin, et al.
Published: (2024)
by: Nilhan Töyer Şahin, et al.
Published: (2024)
Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards
by: Huang, Yangsibo, et al.
Published: (2025)
by: Huang, Yangsibo, et al.
Published: (2025)
Dynamic Power Management Dataset
by: Akbari, Milad
Published: (2025)
by: Akbari, Milad
Published: (2025)
Near Exact Privacy Amplification for Matrix Mechanisms
by: Choquette-Choo, Christopher A., et al.
Published: (2024)
by: Choquette-Choo, Christopher A., et al.
Published: (2024)
Lessons from Defending Gemini Against Indirect Prompt Injections
by: Shi, Chongyang, et al.
Published: (2025)
by: Shi, Chongyang, et al.
Published: (2025)
Institute of marine science of the university of Miami / Robert L. Beutel
by: Beutel Robert, L
Published: (1963)
by: Beutel Robert, L
Published: (1963)
Knowledge Distillation Using Frontier Open-source LLMs: Generalizability and the Role of Synthetic Data
by: Shirgaonkar, Anup, et al.
Published: (2024)
by: Shirgaonkar, Anup, et al.
Published: (2024)
Deliberative Alignment: Reasoning Enables Safer Language Models
by: Guan, Melody Y., et al.
Published: (2024)
by: Guan, Melody Y., et al.
Published: (2024)
The Multiparameter Frontier: Metrological Hierarchy and Robustness in Dispersive Quantum Interferometry
by: de Moura, Lucas Ferreira R., et al.
Published: (2026)
by: de Moura, Lucas Ferreira R., et al.
Published: (2026)
Synthetic Query Generation for Privacy-Preserving Deep Retrieval Systems using Differentially Private Language Models
by: Carranza, Aldo Gael, et al.
Published: (2023)
by: Carranza, Aldo Gael, et al.
Published: (2023)
Half Heusler alloy CoVSn as self-supported electrocatalyst for hydrogen evolution reaction
by: Gujjar, Deepak, et al.
Published: (2024)
by: Gujjar, Deepak, et al.
Published: (2024)
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
by: Wu, Tong, et al.
Published: (2024)
by: Wu, Tong, et al.
Published: (2024)
Similar Items
-
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
by: Wallace, Eric, et al.
Published: (2024) -
Auditing Private Prediction
by: Chadha, Karan, et al.
Published: (2024) -
Privacy Auditing of Large Language Models
by: Panda, Ashwinee, et al.
Published: (2025) -
User Inference Attacks on Large Language Models
by: Kandpal, Nikhil, et al.
Published: (2023) -
LLMs unlock new paths to monetizing exploits
by: Carlini, Nicholas, et al.
Published: (2025)