Optimization-based Prompt Injection Attack to LLM-as-a-Judge
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Jiawen, Yuan, Zenghui, Liu, Yinuo, Huang, Yue, Zhou, Pan, Sun, Lichao, Gong, Neil Zhenqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Prompt Injection Attack to Tool Selection in LLM Agents
von: Shi, Jiawen, et al.
Veröffentlicht: (2025)
von: Shi, Jiawen, et al.
Veröffentlicht: (2025)
BadToken: Token-level Backdoor Attacks to Multi-modal Large Language Models
von: Yuan, Zenghui, et al.
Veröffentlicht: (2025)
von: Yuan, Zenghui, et al.
Veröffentlicht: (2025)
Poisoned-MRAG: Knowledge Poisoning Attacks to Multimodal Retrieval Augmented Generation
von: Liu, Yinuo, et al.
Veröffentlicht: (2025)
von: Liu, Yinuo, et al.
Veröffentlicht: (2025)
ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data
von: Wang, Reachal, et al.
Veröffentlicht: (2025)
von: Wang, Reachal, et al.
Veröffentlicht: (2025)
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
von: Liu, Yinuo, et al.
Veröffentlicht: (2025)
von: Liu, Yinuo, et al.
Veröffentlicht: (2025)
DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models
von: Yuan, Zenghui, et al.
Veröffentlicht: (2025)
von: Yuan, Zenghui, et al.
Veröffentlicht: (2025)
A Critical Evaluation of Defenses against Prompt Injection Attacks
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
von: Shao, Zedian, et al.
Veröffentlicht: (2024)
von: Shao, Zedian, et al.
Veröffentlicht: (2024)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
von: Liu, Yupei, et al.
Veröffentlicht: (2023)
von: Liu, Yupei, et al.
Veröffentlicht: (2023)
WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
von: Wang, Xilong, et al.
Veröffentlicht: (2026)
von: Wang, Xilong, et al.
Veröffentlicht: (2026)
BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
von: Tie, Guiyao, et al.
Veröffentlicht: (2026)
von: Tie, Guiyao, et al.
Veröffentlicht: (2026)
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
von: Zhang, Mohan, et al.
Veröffentlicht: (2026)
von: Zhang, Mohan, et al.
Veröffentlicht: (2026)
PromptLocate: Localizing Prompt Injection Attacks
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
SecInfer: Preventing Prompt Injection via Inference-time Scaling
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
von: Wang, Zhilong, et al.
Veröffentlicht: (2025)
von: Wang, Zhilong, et al.
Veröffentlicht: (2025)
Virtual Context: Enhancing Jailbreak Attacks with Special Token Injection
von: Zhou, Yuqi, et al.
Veröffentlicht: (2024)
von: Zhou, Yuqi, et al.
Veröffentlicht: (2024)
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
von: Wu, Yuanwei, et al.
Veröffentlicht: (2023)
von: Wu, Yuanwei, et al.
Veröffentlicht: (2023)
VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents
von: Cao, Tri, et al.
Veröffentlicht: (2025)
von: Cao, Tri, et al.
Veröffentlicht: (2025)
PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features
von: Zou, Wei, et al.
Veröffentlicht: (2025)
von: Zou, Wei, et al.
Veröffentlicht: (2025)
Leave My Images Alone: Preventing Multi-Modal Large Language Models from Analyzing Images via Visual Prompt Injection
von: Shao, Zedian, et al.
Veröffentlicht: (2026)
von: Shao, Zedian, et al.
Veröffentlicht: (2026)
Evaluating LLM-based Personal Information Extraction and Countermeasures
von: Liu, Yupei, et al.
Veröffentlicht: (2024)
von: Liu, Yupei, et al.
Veröffentlicht: (2024)
AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks
von: Jia, Yuqi, et al.
Veröffentlicht: (2026)
von: Jia, Yuqi, et al.
Veröffentlicht: (2026)
A Survey on Model Extraction Attacks and Defenses for Large Language Models
von: Zhao, Kaixiang, et al.
Veröffentlicht: (2025)
von: Zhao, Kaixiang, et al.
Veröffentlicht: (2025)
A Survey of Model Extraction Attacks and Defenses in Distributed Computing Environments
von: Zhao, Kaixiang, et al.
Veröffentlicht: (2025)
von: Zhao, Kaixiang, et al.
Veröffentlicht: (2025)
Signed-Prompt: A New Approach to Prevent Prompt Injection Attacks Against LLM-Integrated Applications
von: Suo, Xuchen
Veröffentlicht: (2024)
von: Suo, Xuchen
Veröffentlicht: (2024)
AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency Lens
von: Lu, Lin, et al.
Veröffentlicht: (2024)
von: Lu, Lin, et al.
Veröffentlicht: (2024)
A Systematic Survey of Model Extraction Attacks and Defenses: State-of-the-Art and Perspectives
von: Zhao, Kaixiang, et al.
Veröffentlicht: (2025)
von: Zhao, Kaixiang, et al.
Veröffentlicht: (2025)
Model Poisoning Attacks to Federated Learning via Multi-Round Consistency
von: Xie, Yueqi, et al.
Veröffentlicht: (2024)
von: Xie, Yueqi, et al.
Veröffentlicht: (2024)
PromptArmor: Simple yet Effective Prompt Injection Defenses
von: Shi, Tianneng, et al.
Veröffentlicht: (2025)
von: Shi, Tianneng, et al.
Veröffentlicht: (2025)
Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks
von: Chen, Sizhe, et al.
Veröffentlicht: (2025)
von: Chen, Sizhe, et al.
Veröffentlicht: (2025)
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
von: Maloyan, Narek, et al.
Veröffentlicht: (2025)
von: Maloyan, Narek, et al.
Veröffentlicht: (2025)
Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree
von: Johnson, Sam, et al.
Veröffentlicht: (2025)
von: Johnson, Sam, et al.
Veröffentlicht: (2025)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
von: Jaiswal, Piyush, et al.
Veröffentlicht: (2026)
von: Jaiswal, Piyush, et al.
Veröffentlicht: (2026)
Securing AI Agents Against Prompt Injection Attacks
von: Ramakrishnan, Badrinath, et al.
Veröffentlicht: (2025)
von: Ramakrishnan, Badrinath, et al.
Veröffentlicht: (2025)
Can Indirect Prompt Injection Attacks Be Detected and Removed?
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
PIDP-Attack: Combining Prompt Injection with Database Poisoning Attacks on Retrieval-Augmented Generation Systems
von: Wang, Haozhen, et al.
Veröffentlicht: (2026)
von: Wang, Haozhen, et al.
Veröffentlicht: (2026)
AdapTools: Adaptive Tool-based Indirect Prompt Injection Attacks on Agentic LLMs
von: Wang, Che, et al.
Veröffentlicht: (2026)
von: Wang, Che, et al.
Veröffentlicht: (2026)
Token-Efficient Prompt Injection Attack: Provoking Cessation in LLM Reasoning via Adaptive Token Compression
von: Cui, Yu, et al.
Veröffentlicht: (2025)
von: Cui, Yu, et al.
Veröffentlicht: (2025)
The Vulnerability of LLM Rankers to Prompt Injection Attacks
von: Yin, Yu, et al.
Veröffentlicht: (2026)
von: Yin, Yu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Prompt Injection Attack to Tool Selection in LLM Agents
von: Shi, Jiawen, et al.
Veröffentlicht: (2025) -
BadToken: Token-level Backdoor Attacks to Multi-modal Large Language Models
von: Yuan, Zenghui, et al.
Veröffentlicht: (2025) -
Poisoned-MRAG: Knowledge Poisoning Attacks to Multimodal Retrieval Augmented Generation
von: Liu, Yinuo, et al.
Veröffentlicht: (2025) -
ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data
von: Wang, Reachal, et al.
Veröffentlicht: (2025) -
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
von: Liu, Yinuo, et al.
Veröffentlicht: (2025)