PIArena: A Platform for Prompt Injection Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Geng, Runpeng, Yin, Chenlong, Wang, Yanting, Chen, Ying, Jia, Jinyuan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
di: Geng, Runpeng, et al.
Pubblicazione: (2025)
di: Geng, Runpeng, et al.
Pubblicazione: (2025)
PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses
di: Yin, Chenlong, et al.
Pubblicazione: (2026)
di: Yin, Chenlong, et al.
Pubblicazione: (2026)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
di: Liu, Yupei, et al.
Pubblicazione: (2023)
di: Liu, Yupei, et al.
Pubblicazione: (2023)
AttnTrace: Contextual Attribution of Prompt Injection and Knowledge Corruption
di: Wang, Yanting, et al.
Pubblicazione: (2025)
di: Wang, Yanting, et al.
Pubblicazione: (2025)
TracLLM: A Generic Framework for Attributing Long Context LLMs
di: Wang, Yanting, et al.
Pubblicazione: (2025)
di: Wang, Yanting, et al.
Pubblicazione: (2025)
AgentWatcher: A Rule-based Prompt Injection Monitor
di: Wang, Yanting, et al.
Pubblicazione: (2026)
di: Wang, Yanting, et al.
Pubblicazione: (2026)
UniC-RAG: Universal Knowledge Corruption Attacks to Retrieval-Augmented Generation
di: Geng, Runpeng, et al.
Pubblicazione: (2025)
di: Geng, Runpeng, et al.
Pubblicazione: (2025)
FlashRT: Towards Computationally and Memory Efficient Red-Teaming for Prompt Injection and Knowledge Corruption
di: Wang, Yanting, et al.
Pubblicazione: (2026)
di: Wang, Yanting, et al.
Pubblicazione: (2026)
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
di: Zhang, Mohan, et al.
Pubblicazione: (2026)
di: Zhang, Mohan, et al.
Pubblicazione: (2026)
The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
di: Jia, Feiran, et al.
Pubblicazione: (2024)
di: Jia, Feiran, et al.
Pubblicazione: (2024)
Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures
di: Benjamin, Victoria, et al.
Pubblicazione: (2024)
di: Benjamin, Victoria, et al.
Pubblicazione: (2024)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
di: Shao, Zedian, et al.
Pubblicazione: (2024)
di: Shao, Zedian, et al.
Pubblicazione: (2024)
Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
di: Yang, Xiaoxue, et al.
Pubblicazione: (2025)
di: Yang, Xiaoxue, et al.
Pubblicazione: (2025)
PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features
di: Zou, Wei, et al.
Pubblicazione: (2025)
di: Zou, Wei, et al.
Pubblicazione: (2025)
A Systematic Literature Review on LLM Defenses Against Prompt Injection and Jailbreaking: Expanding NIST Taxonomy
di: Correia, Pedro H. Barcha, et al.
Pubblicazione: (2026)
di: Correia, Pedro H. Barcha, et al.
Pubblicazione: (2026)
SecInfer: Preventing Prompt Injection via Inference-time Scaling
di: Liu, Yupei, et al.
Pubblicazione: (2025)
di: Liu, Yupei, et al.
Pubblicazione: (2025)
AgentArmor: Enforcing Program Analysis on Agent Runtime Trace to Defend Against Prompt Injection
di: Wang, Peiran, et al.
Pubblicazione: (2025)
di: Wang, Peiran, et al.
Pubblicazione: (2025)
Jailbreaking with Universal Multi-Prompts
di: Hsu, Yu-Ling, et al.
Pubblicazione: (2025)
di: Hsu, Yu-Ling, et al.
Pubblicazione: (2025)
PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models
di: Zou, Wei, et al.
Pubblicazione: (2024)
di: Zou, Wei, et al.
Pubblicazione: (2024)
EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
di: Liao, Zeyi, et al.
Pubblicazione: (2024)
di: Liao, Zeyi, et al.
Pubblicazione: (2024)
Fight Back Against Jailbreaking via Prompt Adversarial Tuning
di: Mo, Yichuan, et al.
Pubblicazione: (2024)
di: Mo, Yichuan, et al.
Pubblicazione: (2024)
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
di: Liu, Yinuo, et al.
Pubblicazione: (2025)
di: Liu, Yinuo, et al.
Pubblicazione: (2025)
Prompt2Fingerprint: Plug-and-Play LLM Fingerprinting via Text-to-Weight Generation
di: Chen, Sixu, et al.
Pubblicazione: (2026)
di: Chen, Sixu, et al.
Pubblicazione: (2026)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
di: Saiem, Bijoy Ahmed, et al.
Pubblicazione: (2024)
di: Saiem, Bijoy Ahmed, et al.
Pubblicazione: (2024)
Unlocking Memorization in Large Language Models with Dynamic Soft Prompting
di: Wang, Zhepeng, et al.
Pubblicazione: (2024)
di: Wang, Zhepeng, et al.
Pubblicazione: (2024)
Query-Based Adversarial Prompt Generation
di: Hayase, Jonathan, et al.
Pubblicazione: (2024)
di: Hayase, Jonathan, et al.
Pubblicazione: (2024)
PromptLocate: Localizing Prompt Injection Attacks
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
A Critical Evaluation of Defenses against Prompt Injection Attacks
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
Stealing User Prompts from Mixture of Experts
di: Yona, Itay, et al.
Pubblicazione: (2024)
di: Yona, Itay, et al.
Pubblicazione: (2024)
Certifying LLM Safety against Adversarial Prompting
di: Kumar, Aounon, et al.
Pubblicazione: (2023)
di: Kumar, Aounon, et al.
Pubblicazione: (2023)
RECAP: A Resource-Efficient Method for Adversarial Prompting in Large Language Models
di: Chugh, Rishit
Pubblicazione: (2026)
di: Chugh, Rishit
Pubblicazione: (2026)
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
di: Liang, Buyun, et al.
Pubblicazione: (2025)
di: Liang, Buyun, et al.
Pubblicazione: (2025)
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
di: Paulus, Anselm, et al.
Pubblicazione: (2024)
di: Paulus, Anselm, et al.
Pubblicazione: (2024)
MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?
di: Wahed, Muntasir, et al.
Pubblicazione: (2025)
di: Wahed, Muntasir, et al.
Pubblicazione: (2025)
Prompt Injection as Role Confusion
di: Ye, Charles, et al.
Pubblicazione: (2026)
di: Ye, Charles, et al.
Pubblicazione: (2026)
TRAP: Targeted Random Adversarial Prompt Honeypot for Black-Box Identification
di: Gubri, Martin, et al.
Pubblicazione: (2024)
di: Gubri, Martin, et al.
Pubblicazione: (2024)
Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers
di: Zhao, Andrew, et al.
Pubblicazione: (2025)
di: Zhao, Andrew, et al.
Pubblicazione: (2025)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
di: Reddy, Aashray, et al.
Pubblicazione: (2025)
di: Reddy, Aashray, et al.
Pubblicazione: (2025)
Jatmo: Prompt Injection Defense by Task-Specific Finetuning
di: Piet, Julien, et al.
Pubblicazione: (2023)
di: Piet, Julien, et al.
Pubblicazione: (2023)
LLM Platform Security: Applying a Systematic Evaluation Framework to OpenAI's ChatGPT Plugins
di: Iqbal, Umar, et al.
Pubblicazione: (2023)
di: Iqbal, Umar, et al.
Pubblicazione: (2023)
Documenti analoghi
-
PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
di: Geng, Runpeng, et al.
Pubblicazione: (2025) -
PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses
di: Yin, Chenlong, et al.
Pubblicazione: (2026) -
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
di: Liu, Yupei, et al.
Pubblicazione: (2023) -
AttnTrace: Contextual Attribution of Prompt Injection and Knowledge Corruption
di: Wang, Yanting, et al.
Pubblicazione: (2025) -
TracLLM: A Generic Framework for Attributing Long Context LLMs
di: Wang, Yanting, et al.
Pubblicazione: (2025)