FlashRT: Towards Computationally and Memory Efficient Red-Teaming for Prompt Injection and Knowledge Corruption
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yanting, Yin, Chenlong, Chen, Ying, Jia, Jinyuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses
von: Yin, Chenlong, et al.
Veröffentlicht: (2026)
von: Yin, Chenlong, et al.
Veröffentlicht: (2026)
AttnTrace: Contextual Attribution of Prompt Injection and Knowledge Corruption
von: Wang, Yanting, et al.
Veröffentlicht: (2025)
von: Wang, Yanting, et al.
Veröffentlicht: (2025)
PIArena: A Platform for Prompt Injection Evaluation
von: Geng, Runpeng, et al.
Veröffentlicht: (2026)
von: Geng, Runpeng, et al.
Veröffentlicht: (2026)
PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
von: Geng, Runpeng, et al.
Veröffentlicht: (2025)
von: Geng, Runpeng, et al.
Veröffentlicht: (2025)
AgentWatcher: A Rule-based Prompt Injection Monitor
von: Wang, Yanting, et al.
Veröffentlicht: (2026)
von: Wang, Yanting, et al.
Veröffentlicht: (2026)
UniC-RAG: Universal Knowledge Corruption Attacks to Retrieval-Augmented Generation
von: Geng, Runpeng, et al.
Veröffentlicht: (2025)
von: Geng, Runpeng, et al.
Veröffentlicht: (2025)
PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features
von: Zou, Wei, et al.
Veröffentlicht: (2025)
von: Zou, Wei, et al.
Veröffentlicht: (2025)
SecInfer: Preventing Prompt Injection via Inference-time Scaling
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
EnsembleSHAP: Faithful and Certifiably Robust Attribution for Random Subspace Method
von: Wang, Yanting, et al.
Veröffentlicht: (2026)
von: Wang, Yanting, et al.
Veröffentlicht: (2026)
PromptLocate: Localizing Prompt Injection Attacks
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
FCert: Certifiably Robust Few-Shot Classification in the Era of Foundation Models
von: Wang, Yanting, et al.
Veröffentlicht: (2024)
von: Wang, Yanting, et al.
Veröffentlicht: (2024)
PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models
von: Zou, Wei, et al.
Veröffentlicht: (2024)
von: Zou, Wei, et al.
Veröffentlicht: (2024)
TrojanDec: Data-free Detection of Trojan Inputs in Self-supervised Learning
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
TASO: Jailbreak LLMs via Alternative Template and Suffix Optimization
von: Wang, Yanting, et al.
Veröffentlicht: (2025)
von: Wang, Yanting, et al.
Veröffentlicht: (2025)
DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
A Critical Evaluation of Defenses against Prompt Injection Attacks
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
Whispers of Wealth: Red-Teaming Google's Agent Payments Protocol via Prompt Injection
von: Debi, Tanusree, et al.
Veröffentlicht: (2026)
von: Debi, Tanusree, et al.
Veröffentlicht: (2026)
OpenRT: An Open-Source Red Teaming Framework for Multimodal LLMs
von: Wang, Xin, et al.
Veröffentlicht: (2026)
von: Wang, Xin, et al.
Veröffentlicht: (2026)
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
von: Syros, Georgios, et al.
Veröffentlicht: (2026)
von: Syros, Georgios, et al.
Veröffentlicht: (2026)
Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
von: Pathade, Chetan
Veröffentlicht: (2025)
von: Pathade, Chetan
Veröffentlicht: (2025)
IPI-proxy: An Intercepting Proxy for Red-Teaming Web-Browsing AI Agents Against Indirect Prompt Injection
von: Chia-Pei, et al.
Veröffentlicht: (2026)
von: Chia-Pei, et al.
Veröffentlicht: (2026)
The Vulnerability of LLM Rankers to Prompt Injection Attacks
von: Yin, Yu, et al.
Veröffentlicht: (2026)
von: Yin, Yu, et al.
Veröffentlicht: (2026)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
von: Liu, Yupei, et al.
Veröffentlicht: (2023)
von: Liu, Yupei, et al.
Veröffentlicht: (2023)
ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data
von: Wang, Reachal, et al.
Veröffentlicht: (2025)
von: Wang, Reachal, et al.
Veröffentlicht: (2025)
Trojan Horses in Recruiting: A Red-Teaming Case Study on Indirect Prompt Injection in Standard vs. Reasoning Models
von: Wirth, Manuel
Veröffentlicht: (2026)
von: Wirth, Manuel
Veröffentlicht: (2026)
Persona-Conditioned Adversarial Prompting (PCAP): Multi-Identity Red-Teaming for Enhanced Adversarial Prompt Discovery
von: Morasso, Cristian, et al.
Veröffentlicht: (2026)
von: Morasso, Cristian, et al.
Veröffentlicht: (2026)
AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks
von: Jia, Yuqi, et al.
Veröffentlicht: (2026)
von: Jia, Yuqi, et al.
Veröffentlicht: (2026)
TracLLM: A Generic Framework for Attributing Long Context LLMs
von: Wang, Yanting, et al.
Veröffentlicht: (2025)
von: Wang, Yanting, et al.
Veröffentlicht: (2025)
MMCert: Provable Defense against Adversarial Attacks to Multi-modal Models
von: Wang, Yanting, et al.
Veröffentlicht: (2024)
von: Wang, Yanting, et al.
Veröffentlicht: (2024)
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
von: Kumar, Anurakt, et al.
Veröffentlicht: (2024)
von: Kumar, Anurakt, et al.
Veröffentlicht: (2024)
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
von: Liu, Yanjiang, et al.
Veröffentlicht: (2025)
von: Liu, Yanjiang, et al.
Veröffentlicht: (2025)
Prompt Optimization and Evaluation for LLM Automated Red Teaming
von: Freenor, Michael, et al.
Veröffentlicht: (2025)
von: Freenor, Michael, et al.
Veröffentlicht: (2025)
Defending Against Prompt Injection with DataFilter
von: Wang, Yizhu, et al.
Veröffentlicht: (2025)
von: Wang, Yizhu, et al.
Veröffentlicht: (2025)
Red Teaming Large Reasoning Models
von: Chen, Jiawei, et al.
Veröffentlicht: (2025)
von: Chen, Jiawei, et al.
Veröffentlicht: (2025)
PromptSleuth: Detecting Prompt Injection via Semantic Intent Invariance
von: Wang, Mengxiao, et al.
Veröffentlicht: (2025)
von: Wang, Mengxiao, et al.
Veröffentlicht: (2025)
A Red Teaming Roadmap Towards System-Level Safety
von: Wang, Zifan, et al.
Veröffentlicht: (2025)
von: Wang, Zifan, et al.
Veröffentlicht: (2025)
Red Teaming Methodology for Design Obfuscation
von: Liu, Yuntao, et al.
Veröffentlicht: (2025)
von: Liu, Yuntao, et al.
Veröffentlicht: (2025)
Defending Against Prompt Injection With a Few DefensiveTokens
von: Chen, Sizhe, et al.
Veröffentlicht: (2025)
von: Chen, Sizhe, et al.
Veröffentlicht: (2025)
PromptShield: Deployable Detection for Prompt Injection Attacks
von: Jacob, Dennis, et al.
Veröffentlicht: (2025)
von: Jacob, Dennis, et al.
Veröffentlicht: (2025)
AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
von: Ying, Zonghao, et al.
Veröffentlicht: (2026)
von: Ying, Zonghao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses
von: Yin, Chenlong, et al.
Veröffentlicht: (2026) -
AttnTrace: Contextual Attribution of Prompt Injection and Knowledge Corruption
von: Wang, Yanting, et al.
Veröffentlicht: (2025) -
PIArena: A Platform for Prompt Injection Evaluation
von: Geng, Runpeng, et al.
Veröffentlicht: (2026) -
PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
von: Geng, Runpeng, et al.
Veröffentlicht: (2025) -
AgentWatcher: A Rule-based Prompt Injection Monitor
von: Wang, Yanting, et al.
Veröffentlicht: (2026)