Optimization-based Prompt Injection Attack to LLM-as-a-Judge
Fuente:
arXiv
Salvato in:
| Autori principali: | Shi, Jiawen, Yuan, Zenghui, Liu, Yinuo, Huang, Yue, Zhou, Pan, Sun, Lichao, Gong, Neil Zhenqiang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Prompt Injection Attack to Tool Selection in LLM Agents
di: Shi, Jiawen, et al.
Pubblicazione: (2025)
di: Shi, Jiawen, et al.
Pubblicazione: (2025)
BadToken: Token-level Backdoor Attacks to Multi-modal Large Language Models
di: Yuan, Zenghui, et al.
Pubblicazione: (2025)
di: Yuan, Zenghui, et al.
Pubblicazione: (2025)
Poisoned-MRAG: Knowledge Poisoning Attacks to Multimodal Retrieval Augmented Generation
di: Liu, Yinuo, et al.
Pubblicazione: (2025)
di: Liu, Yinuo, et al.
Pubblicazione: (2025)
ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data
di: Wang, Reachal, et al.
Pubblicazione: (2025)
di: Wang, Reachal, et al.
Pubblicazione: (2025)
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
di: Liu, Yinuo, et al.
Pubblicazione: (2025)
di: Liu, Yinuo, et al.
Pubblicazione: (2025)
DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
di: Liu, Yupei, et al.
Pubblicazione: (2025)
di: Liu, Yupei, et al.
Pubblicazione: (2025)
Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models
di: Yuan, Zenghui, et al.
Pubblicazione: (2025)
di: Yuan, Zenghui, et al.
Pubblicazione: (2025)
A Critical Evaluation of Defenses against Prompt Injection Attacks
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
di: Shao, Zedian, et al.
Pubblicazione: (2024)
di: Shao, Zedian, et al.
Pubblicazione: (2024)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
di: Liu, Yupei, et al.
Pubblicazione: (2023)
di: Liu, Yupei, et al.
Pubblicazione: (2023)
WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
di: Wang, Xilong, et al.
Pubblicazione: (2026)
di: Wang, Xilong, et al.
Pubblicazione: (2026)
BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
di: Tie, Guiyao, et al.
Pubblicazione: (2026)
di: Tie, Guiyao, et al.
Pubblicazione: (2026)
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
di: Zhang, Mohan, et al.
Pubblicazione: (2026)
di: Zhang, Mohan, et al.
Pubblicazione: (2026)
PromptLocate: Localizing Prompt Injection Attacks
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
SecInfer: Preventing Prompt Injection via Inference-time Scaling
di: Liu, Yupei, et al.
Pubblicazione: (2025)
di: Liu, Yupei, et al.
Pubblicazione: (2025)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
di: Wang, Zhilong, et al.
Pubblicazione: (2025)
di: Wang, Zhilong, et al.
Pubblicazione: (2025)
Virtual Context: Enhancing Jailbreak Attacks with Special Token Injection
di: Zhou, Yuqi, et al.
Pubblicazione: (2024)
di: Zhou, Yuqi, et al.
Pubblicazione: (2024)
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
di: Wu, Yuanwei, et al.
Pubblicazione: (2023)
di: Wu, Yuanwei, et al.
Pubblicazione: (2023)
VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents
di: Cao, Tri, et al.
Pubblicazione: (2025)
di: Cao, Tri, et al.
Pubblicazione: (2025)
PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features
di: Zou, Wei, et al.
Pubblicazione: (2025)
di: Zou, Wei, et al.
Pubblicazione: (2025)
Leave My Images Alone: Preventing Multi-Modal Large Language Models from Analyzing Images via Visual Prompt Injection
di: Shao, Zedian, et al.
Pubblicazione: (2026)
di: Shao, Zedian, et al.
Pubblicazione: (2026)
Evaluating LLM-based Personal Information Extraction and Countermeasures
di: Liu, Yupei, et al.
Pubblicazione: (2024)
di: Liu, Yupei, et al.
Pubblicazione: (2024)
AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks
di: Jia, Yuqi, et al.
Pubblicazione: (2026)
di: Jia, Yuqi, et al.
Pubblicazione: (2026)
A Survey on Model Extraction Attacks and Defenses for Large Language Models
di: Zhao, Kaixiang, et al.
Pubblicazione: (2025)
di: Zhao, Kaixiang, et al.
Pubblicazione: (2025)
A Survey of Model Extraction Attacks and Defenses in Distributed Computing Environments
di: Zhao, Kaixiang, et al.
Pubblicazione: (2025)
di: Zhao, Kaixiang, et al.
Pubblicazione: (2025)
Signed-Prompt: A New Approach to Prevent Prompt Injection Attacks Against LLM-Integrated Applications
di: Suo, Xuchen
Pubblicazione: (2024)
di: Suo, Xuchen
Pubblicazione: (2024)
AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency Lens
di: Lu, Lin, et al.
Pubblicazione: (2024)
di: Lu, Lin, et al.
Pubblicazione: (2024)
A Systematic Survey of Model Extraction Attacks and Defenses: State-of-the-Art and Perspectives
di: Zhao, Kaixiang, et al.
Pubblicazione: (2025)
di: Zhao, Kaixiang, et al.
Pubblicazione: (2025)
Model Poisoning Attacks to Federated Learning via Multi-Round Consistency
di: Xie, Yueqi, et al.
Pubblicazione: (2024)
di: Xie, Yueqi, et al.
Pubblicazione: (2024)
PromptArmor: Simple yet Effective Prompt Injection Defenses
di: Shi, Tianneng, et al.
Pubblicazione: (2025)
di: Shi, Tianneng, et al.
Pubblicazione: (2025)
Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks
di: Chen, Sizhe, et al.
Pubblicazione: (2025)
di: Chen, Sizhe, et al.
Pubblicazione: (2025)
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
di: Maloyan, Narek, et al.
Pubblicazione: (2025)
di: Maloyan, Narek, et al.
Pubblicazione: (2025)
Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree
di: Johnson, Sam, et al.
Pubblicazione: (2025)
di: Johnson, Sam, et al.
Pubblicazione: (2025)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
di: Jaiswal, Piyush, et al.
Pubblicazione: (2026)
di: Jaiswal, Piyush, et al.
Pubblicazione: (2026)
Securing AI Agents Against Prompt Injection Attacks
di: Ramakrishnan, Badrinath, et al.
Pubblicazione: (2025)
di: Ramakrishnan, Badrinath, et al.
Pubblicazione: (2025)
Can Indirect Prompt Injection Attacks Be Detected and Removed?
di: Chen, Yulin, et al.
Pubblicazione: (2025)
di: Chen, Yulin, et al.
Pubblicazione: (2025)
PIDP-Attack: Combining Prompt Injection with Database Poisoning Attacks on Retrieval-Augmented Generation Systems
di: Wang, Haozhen, et al.
Pubblicazione: (2026)
di: Wang, Haozhen, et al.
Pubblicazione: (2026)
AdapTools: Adaptive Tool-based Indirect Prompt Injection Attacks on Agentic LLMs
di: Wang, Che, et al.
Pubblicazione: (2026)
di: Wang, Che, et al.
Pubblicazione: (2026)
Token-Efficient Prompt Injection Attack: Provoking Cessation in LLM Reasoning via Adaptive Token Compression
di: Cui, Yu, et al.
Pubblicazione: (2025)
di: Cui, Yu, et al.
Pubblicazione: (2025)
The Vulnerability of LLM Rankers to Prompt Injection Attacks
di: Yin, Yu, et al.
Pubblicazione: (2026)
di: Yin, Yu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Prompt Injection Attack to Tool Selection in LLM Agents
di: Shi, Jiawen, et al.
Pubblicazione: (2025) -
BadToken: Token-level Backdoor Attacks to Multi-modal Large Language Models
di: Yuan, Zenghui, et al.
Pubblicazione: (2025) -
Poisoned-MRAG: Knowledge Poisoning Attacks to Multimodal Retrieval Augmented Generation
di: Liu, Yinuo, et al.
Pubblicazione: (2025) -
ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data
di: Wang, Reachal, et al.
Pubblicazione: (2025) -
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
di: Liu, Yinuo, et al.
Pubblicazione: (2025)