The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wallace, Eric, Xiao, Kai, Leike, Reimar, Weng, Lilian, Heidecke, Johannes, Beutel, Alex |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
par: Guo, Chuan, et autres
Publié: (2026)
par: Guo, Chuan, et autres
Publié: (2026)
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
par: Beutel, Alex, et autres
Publié: (2024)
par: Beutel, Alex, et autres
Publié: (2024)
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
par: Wu, Tong, et autres
Publié: (2024)
par: Wu, Tong, et autres
Publié: (2024)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
par: Xu, Jiashu, et autres
Publié: (2023)
par: Xu, Jiashu, et autres
Publié: (2023)
Instructional Fingerprinting of Large Language Models
par: Xu, Jiashu, et autres
Publié: (2024)
par: Xu, Jiashu, et autres
Publié: (2024)
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
par: Yan, Jun, et autres
Publié: (2023)
par: Yan, Jun, et autres
Publié: (2023)
Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities
par: Geng, Jiahui, et autres
Publié: (2025)
par: Geng, Jiahui, et autres
Publié: (2025)
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
par: Qi, Zhenting, et autres
Publié: (2024)
par: Qi, Zhenting, et autres
Publié: (2024)
Can We Infer Confidential Properties of Training Data from LLMs?
par: Huang, Pengrun, et autres
Publié: (2025)
par: Huang, Pengrun, et autres
Publié: (2025)
Coercing LLMs to do and reveal (almost) anything
par: Geiping, Jonas, et autres
Publié: (2024)
par: Geiping, Jonas, et autres
Publié: (2024)
Adaptive Instruction Composition for Automated LLM Red-Teaming
par: Zymet, Jesse, et autres
Publié: (2026)
par: Zymet, Jesse, et autres
Publié: (2026)
Intrinsic Fingerprint of LLMs: Continue Training is NOT All You Need to Steal A Model!
par: Yoon, Do-hyeon, et autres
Publié: (2025)
par: Yoon, Do-hyeon, et autres
Publié: (2025)
A Study of Backdoors in Instruction Fine-tuned Language Models
par: Raghuram, Jayaram, et autres
Publié: (2024)
par: Raghuram, Jayaram, et autres
Publié: (2024)
Trading Inference-Time Compute for Adversarial Robustness
par: Zaremba, Wojciech, et autres
Publié: (2025)
par: Zaremba, Wojciech, et autres
Publié: (2025)
Cross-Task Defense: Instruction-Tuning LLMs for Content Safety
par: Fu, Yu, et autres
Publié: (2024)
par: Fu, Yu, et autres
Publié: (2024)
MARAGE: Transferable Multi-Model Adversarial Attack for Retrieval-Augmented Generation Data Extraction
par: Hu, Xiao, et autres
Publié: (2025)
par: Hu, Xiao, et autres
Publié: (2025)
Instruction Backdoor Attacks Against Customized LLMs
par: Zhang, Rui, et autres
Publié: (2024)
par: Zhang, Rui, et autres
Publié: (2024)
How Vulnerable Are Edge LLMs?
par: Ding, Ao, et autres
Publié: (2026)
par: Ding, Ao, et autres
Publié: (2026)
UCD: Unlearning in LLMs via Contrastive Decoding
par: Suriyakumar, Vinith M., et autres
Publié: (2025)
par: Suriyakumar, Vinith M., et autres
Publié: (2025)
Sockpuppetting: Jailbreaking LLMs by Combining Prefilling with Optimization
par: Dotsinski, Asen, et autres
Publié: (2026)
par: Dotsinski, Asen, et autres
Publié: (2026)
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
par: Xiong, Alexander, et autres
Publié: (2025)
par: Xiong, Alexander, et autres
Publié: (2025)
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
par: Vega, Jason, et autres
Publié: (2023)
par: Vega, Jason, et autres
Publié: (2023)
Permute-and-Flip: An optimally stable and watermarkable decoder for LLMs
par: Zhao, Xuandong, et autres
Publié: (2024)
par: Zhao, Xuandong, et autres
Publié: (2024)
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
par: Sivapiromrat, Sanhanat, et autres
Publié: (2025)
par: Sivapiromrat, Sanhanat, et autres
Publié: (2025)
Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-Tuning
par: Hasan, Adib, et autres
Publié: (2024)
par: Hasan, Adib, et autres
Publié: (2024)
Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs
par: Mathew, Yohan, et autres
Publié: (2024)
par: Mathew, Yohan, et autres
Publié: (2024)
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
par: Brown, Hannah, et autres
Publié: (2024)
par: Brown, Hannah, et autres
Publié: (2024)
Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs
par: Price, Sara, et autres
Publié: (2024)
par: Price, Sara, et autres
Publié: (2024)
Bias Amplification in RAG: Poisoning Knowledge Retrieval to Steer LLMs
par: Wang, Linlin, et autres
Publié: (2025)
par: Wang, Linlin, et autres
Publié: (2025)
Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models
par: Hu, Kai, et autres
Publié: (2025)
par: Hu, Kai, et autres
Publié: (2025)
Cyber-Zero: Training Cybersecurity Agents without Runtime
par: Zhuo, Terry Yue, et autres
Publié: (2025)
par: Zhuo, Terry Yue, et autres
Publié: (2025)
Randomized Masked Finetuning: An Efficient Way to Mitigate Memorization of PIIs in LLMs
par: Joshi, Kunj, et autres
Publié: (2025)
par: Joshi, Kunj, et autres
Publié: (2025)
Securing Large Language Models (LLMs) from Prompt Injection Attacks
par: Suri, Omar Farooq Khan, et autres
Publié: (2025)
par: Suri, Omar Farooq Khan, et autres
Publié: (2025)
LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
par: Li, Ran, et autres
Publié: (2025)
par: Li, Ran, et autres
Publié: (2025)
SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It)
par: Meeus, Matthieu, et autres
Publié: (2024)
par: Meeus, Matthieu, et autres
Publié: (2024)
Privacy-Preserving Synthetic Review Generation with Diverse Writing Styles Using LLMs
par: Atwal, Tevin, et autres
Publié: (2025)
par: Atwal, Tevin, et autres
Publié: (2025)
When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models
par: Wang, Kai, et autres
Publié: (2025)
par: Wang, Kai, et autres
Publié: (2025)
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
par: Kaneko, Masahiro, et autres
Publié: (2025)
par: Kaneko, Masahiro, et autres
Publié: (2025)
Implicit Identity Technologies for LLMs: Fingerprinting and Watermarking across Datasets, Models, and Generated Content
par: Liu, Bing, et autres
Publié: (2026)
par: Liu, Bing, et autres
Publié: (2026)
SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
par: Xue, Eric, et autres
Publié: (2025)
par: Xue, Eric, et autres
Publié: (2025)
Documents similaires
-
IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
par: Guo, Chuan, et autres
Publié: (2026) -
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
par: Beutel, Alex, et autres
Publié: (2024) -
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
par: Wu, Tong, et autres
Publié: (2024) -
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
par: Xu, Jiashu, et autres
Publié: (2023) -
Instructional Fingerprinting of Large Language Models
par: Xu, Jiashu, et autres
Publié: (2024)