Know Thy Enemy: Securing LLMs Against Prompt Injection via Diverse Data Synthesis and Instruction-Level Chain-of-Thought Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chang, Zhiyuan, Li, Mingyang, Huang, Yuekai, Jiang, Ziyou, Jia, Xiaojun, Xiong, Qian, Wang, Junjie, Li, Zhaoyang, Wang, Qing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
All Changes May Have Invariant Principles: Improving Ever-Shifting Harmful Meme Detection via Design Concept Reproduction
von: Jiang, Ziyou, et al.
Veröffentlicht: (2026)
von: Jiang, Ziyou, et al.
Veröffentlicht: (2026)
One Shot Dominance: Knowledge Poisoning Attack on Retrieval-Augmented Generation Systems
von: Chang, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Chang, Zhiyuan, et al.
Veröffentlicht: (2025)
What External Knowledge is Preferred by LLMs? Characterizing and Exploring Chain of Evidence in Imperfect Context for Multi-Hop QA
von: Chang, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Chang, Zhiyuan, et al.
Veröffentlicht: (2024)
Mimicking the Familiar: Dynamic Command Generation for Information Theft Attacks in LLM Tool-Learning System
von: Jiang, Ziyou, et al.
Veröffentlicht: (2025)
von: Jiang, Ziyou, et al.
Veröffentlicht: (2025)
Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems
von: Xiong, Qian, et al.
Veröffentlicht: (2025)
von: Xiong, Qian, et al.
Veröffentlicht: (2025)
Fall into a Pit, Gain in a Wit: Cognitive-Guided Harmful Meme Detection via Misjudgment Risk Pattern Retrieval
von: Wang, Wenshuo, et al.
Veröffentlicht: (2025)
von: Wang, Wenshuo, et al.
Veröffentlicht: (2025)
From Allies to Adversaries: Manipulating LLM Tool-Calling through Adversarial Injection
von: Wang, Haowei, et al.
Veröffentlicht: (2024)
von: Wang, Haowei, et al.
Veröffentlicht: (2024)
Emerging from Ground: Addressing Intent Deviation in Tool-Using Agents via Deriving Real Calls into Virtual Trajectories
von: Xiong, Qian, et al.
Veröffentlicht: (2026)
von: Xiong, Qian, et al.
Veröffentlicht: (2026)
Are All Prompt Components Value-Neutral? Understanding the Heterogeneous Adversarial Robustness of Dissected Prompt in Large Language Models
von: Zheng, Yujia, et al.
Veröffentlicht: (2025)
von: Zheng, Yujia, et al.
Veröffentlicht: (2025)
Adversarial Robustness of Open-source Text Classification Models and Fine-Tuning Chains
von: Qin, Hao, et al.
Veröffentlicht: (2024)
von: Qin, Hao, et al.
Veröffentlicht: (2024)
Joint-GCG: Unified Gradient-Based Poisoning Attacks on Retrieval-Augmented Generation Systems
von: Wang, Haowei, et al.
Veröffentlicht: (2025)
von: Wang, Haowei, et al.
Veröffentlicht: (2025)
Controllable Navigation Instruction Generation with Chain of Thought Prompting
von: Kong, Xianghao, et al.
Veröffentlicht: (2024)
von: Kong, Xianghao, et al.
Veröffentlicht: (2024)
VEglue: Testing Visual Entailment Systems via Object-Aligned Joint Erasing
von: Chang, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Chang, Zhiyuan, et al.
Veröffentlicht: (2024)
Securing AI Agents Against Prompt Injection Attacks
von: Ramakrishnan, Badrinath, et al.
Veröffentlicht: (2025)
von: Ramakrishnan, Badrinath, et al.
Veröffentlicht: (2025)
Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs
von: Wang, Jiawen, et al.
Veröffentlicht: (2025)
von: Wang, Jiawen, et al.
Veröffentlicht: (2025)
CoT-Drive: Efficient Motion Forecasting for Autonomous Driving with LLMs and Chain-of-Thought Prompting
von: Liao, Haicheng, et al.
Veröffentlicht: (2025)
von: Liao, Haicheng, et al.
Veröffentlicht: (2025)
VulRTex: A Reasoning-Guided Approach to Identify Vulnerabilities from Rich-Text Issue Report
von: Jiang, Ziyou, et al.
Veröffentlicht: (2025)
von: Jiang, Ziyou, et al.
Veröffentlicht: (2025)
Learning to Edit Knowledge via Instruction-based Chain-of-Thought Prompting
von: Fu, Jinhu, et al.
Veröffentlicht: (2026)
von: Fu, Jinhu, et al.
Veröffentlicht: (2026)
ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior
von: Lu, Weikai, et al.
Veröffentlicht: (2025)
von: Lu, Weikai, et al.
Veröffentlicht: (2025)
Copyright Infringement Risk Reduction via Chain-of-Thought and Task Instruction Prompting
von: Sarna, Neeraj, et al.
Veröffentlicht: (2025)
von: Sarna, Neeraj, et al.
Veröffentlicht: (2025)
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
von: Xiang, Zhen, et al.
Veröffentlicht: (2024)
von: Xiang, Zhen, et al.
Veröffentlicht: (2024)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
von: Jaiswal, Piyush, et al.
Veröffentlicht: (2026)
von: Jaiswal, Piyush, et al.
Veröffentlicht: (2026)
LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
von: Li, Chloe, et al.
Veröffentlicht: (2025)
von: Li, Chloe, et al.
Veröffentlicht: (2025)
Secure-Instruct: An Automated Pipeline for Synthesizing Instruction-Tuning Datasets Using LLMs for Secure Code Generation
von: Li, Junjie, et al.
Veröffentlicht: (2025)
von: Li, Junjie, et al.
Veröffentlicht: (2025)
Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks
von: Xiang, Chong, et al.
Veröffentlicht: (2026)
von: Xiang, Chong, et al.
Veröffentlicht: (2026)
Chain-of-Thought Reasoning Without Prompting
von: Wang, Xuezhi, et al.
Veröffentlicht: (2024)
von: Wang, Xuezhi, et al.
Veröffentlicht: (2024)
PromptEnhancer: A Simple Approach to Enhance Text-to-Image Models via Chain-of-Thought Prompt Rewriting
von: Wang, Linqing, et al.
Veröffentlicht: (2025)
von: Wang, Linqing, et al.
Veröffentlicht: (2025)
Play Guessing Game with LLM: Indirect Jailbreak Attack with Implicit Clues
von: Chang, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Chang, Zhiyuan, et al.
Veröffentlicht: (2024)
Repairing Catastrophic-Neglect in Text-to-Image Diffusion Models via Attention-Guided Feature Enhancement
von: Chang, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Chang, Zhiyuan, et al.
Veröffentlicht: (2024)
TokenSkip: Controllable Chain-of-Thought Compression in LLMs
von: Xia, Heming, et al.
Veröffentlicht: (2025)
von: Xia, Heming, et al.
Veröffentlicht: (2025)
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
von: Evtimov, Ivan, et al.
Veröffentlicht: (2025)
von: Evtimov, Ivan, et al.
Veröffentlicht: (2025)
Defending Against Prompt Injection with DataFilter
von: Wang, Yizhu, et al.
Veröffentlicht: (2025)
von: Wang, Yizhu, et al.
Veröffentlicht: (2025)
Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges
von: Eiras, Francisco, et al.
Veröffentlicht: (2025)
von: Eiras, Francisco, et al.
Veröffentlicht: (2025)
Prompt-with-Me: in-IDE Structured Prompt Management for LLM-Driven Software Engineering
von: Li, Ziyou, et al.
Veröffentlicht: (2025)
von: Li, Ziyou, et al.
Veröffentlicht: (2025)
Beyond Chain-of-Thought: A Survey of Chain-of-X Paradigms for LLMs
von: Xia, Yu, et al.
Veröffentlicht: (2024)
von: Xia, Yu, et al.
Veröffentlicht: (2024)
Adversarial Testing for Visual Grounding via Image-Aware Property Reduction
von: Chang, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Chang, Zhiyuan, et al.
Veröffentlicht: (2024)
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
von: Jiang, Qing, et al.
Veröffentlicht: (2025)
von: Jiang, Qing, et al.
Veröffentlicht: (2025)
CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks
von: Wang, Yanan, et al.
Veröffentlicht: (2025)
von: Wang, Yanan, et al.
Veröffentlicht: (2025)
Image-based Prompt Injection: Hijacking Multimodal LLMs through Visually Embedded Adversarial Instructions
von: Nagaraja, Neha, et al.
Veröffentlicht: (2026)
von: Nagaraja, Neha, et al.
Veröffentlicht: (2026)
Keypoint-based Progressive Chain-of-Thought Distillation for LLMs
von: Feng, Kaituo, et al.
Veröffentlicht: (2024)
von: Feng, Kaituo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
All Changes May Have Invariant Principles: Improving Ever-Shifting Harmful Meme Detection via Design Concept Reproduction
von: Jiang, Ziyou, et al.
Veröffentlicht: (2026) -
One Shot Dominance: Knowledge Poisoning Attack on Retrieval-Augmented Generation Systems
von: Chang, Zhiyuan, et al.
Veröffentlicht: (2025) -
What External Knowledge is Preferred by LLMs? Characterizing and Exploring Chain of Evidence in Imperfect Context for Multi-Hop QA
von: Chang, Zhiyuan, et al.
Veröffentlicht: (2024) -
Mimicking the Familiar: Dynamic Command Generation for Information Theft Attacks in LLM Tool-Learning System
von: Jiang, Ziyou, et al.
Veröffentlicht: (2025) -
Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems
von: Xiong, Qian, et al.
Veröffentlicht: (2025)