PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhu, Kaijie, Wang, Jindong, Zhou, Jiaheng, Wang, Zichen, Chen, Hao, Wang, Yidong, Yang, Linyi, Ye, Wei, Zhang, Yue, Gong, Neil Zhenqiang, Xie, Xing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data
di: Wang, Reachal, et al.
Pubblicazione: (2025)
di: Wang, Reachal, et al.
Pubblicazione: (2025)
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
di: Zhu, Kaijie, et al.
Pubblicazione: (2025)
di: Zhu, Kaijie, et al.
Pubblicazione: (2025)
SecInfer: Preventing Prompt Injection via Inference-time Scaling
di: Liu, Yupei, et al.
Pubblicazione: (2025)
di: Liu, Yupei, et al.
Pubblicazione: (2025)
A Critical Evaluation of Defenses against Prompt Injection Attacks
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
Prompt Injection Attack to Tool Selection in LLM Agents
di: Shi, Jiawen, et al.
Pubblicazione: (2025)
di: Shi, Jiawen, et al.
Pubblicazione: (2025)
Provably Robust Federated Reinforcement Learning
di: Fang, Minghong, et al.
Pubblicazione: (2025)
di: Fang, Minghong, et al.
Pubblicazione: (2025)
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
di: Liu, Yinuo, et al.
Pubblicazione: (2025)
di: Liu, Yinuo, et al.
Pubblicazione: (2025)
Optimization-based Prompt Injection Attack to LLM-as-a-Judge
di: Shi, Jiawen, et al.
Pubblicazione: (2024)
di: Shi, Jiawen, et al.
Pubblicazione: (2024)
PromptArmor: Simple yet Effective Prompt Injection Defenses
di: Shi, Tianneng, et al.
Pubblicazione: (2025)
di: Shi, Tianneng, et al.
Pubblicazione: (2025)
DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
di: Liu, Yupei, et al.
Pubblicazione: (2025)
di: Liu, Yupei, et al.
Pubblicazione: (2025)
AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks
di: Jia, Yuqi, et al.
Pubblicazione: (2026)
di: Jia, Yuqi, et al.
Pubblicazione: (2026)
Model Poisoning Attacks to Federated Learning via Multi-Round Consistency
di: Xie, Yueqi, et al.
Pubblicazione: (2024)
di: Xie, Yueqi, et al.
Pubblicazione: (2024)
Refusing Safe Prompts for Multi-modal Large Language Models
di: Shao, Zedian, et al.
Pubblicazione: (2024)
di: Shao, Zedian, et al.
Pubblicazione: (2024)
SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts
di: Xin, Yuan, et al.
Pubblicazione: (2026)
di: Xin, Yuan, et al.
Pubblicazione: (2026)
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
di: Cao, Tri, et al.
Pubblicazione: (2026)
di: Cao, Tri, et al.
Pubblicazione: (2026)
Robustness of Vision Foundation Models to Common Perturbations
di: Liu, Hongbin, et al.
Pubblicazione: (2026)
di: Liu, Hongbin, et al.
Pubblicazione: (2026)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
di: Shao, Zedian, et al.
Pubblicazione: (2024)
di: Shao, Zedian, et al.
Pubblicazione: (2024)
Evaluating LLM-based Personal Information Extraction and Countermeasures
di: Liu, Yupei, et al.
Pubblicazione: (2024)
di: Liu, Yupei, et al.
Pubblicazione: (2024)
LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment Training
di: Gong, Yuyang, et al.
Pubblicazione: (2026)
di: Gong, Yuyang, et al.
Pubblicazione: (2026)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
di: Liu, Yupei, et al.
Pubblicazione: (2023)
di: Liu, Yupei, et al.
Pubblicazione: (2023)
AudioMarkBench: Benchmarking Robustness of Audio Watermarking
di: Liu, Hongbin, et al.
Pubblicazione: (2024)
di: Liu, Hongbin, et al.
Pubblicazione: (2024)
Are All Prompt Components Value-Neutral? Understanding the Heterogeneous Adversarial Robustness of Dissected Prompt in Large Language Models
di: Zheng, Yujia, et al.
Pubblicazione: (2025)
di: Zheng, Yujia, et al.
Pubblicazione: (2025)
PromptLocate: Localizing Prompt Injection Attacks
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
di: Jia, Yuqi, et al.
Pubblicazione: (2025)
Leave My Images Alone: Preventing Multi-Modal Large Language Models from Analyzing Images via Visual Prompt Injection
di: Shao, Zedian, et al.
Pubblicazione: (2026)
di: Shao, Zedian, et al.
Pubblicazione: (2026)
Pruning Graphs by Adversarial Robustness Evaluation to Strengthen GNN Defenses
di: Wang, Yongyu
Pubblicazione: (2025)
di: Wang, Yongyu
Pubblicazione: (2025)
Prompt Inversion Attack against Collaborative Inference of Large Language Models
di: Qu, Wenjie, et al.
Pubblicazione: (2025)
di: Qu, Wenjie, et al.
Pubblicazione: (2025)
BadToken: Token-level Backdoor Attacks to Multi-modal Large Language Models
di: Yuan, Zenghui, et al.
Pubblicazione: (2025)
di: Yuan, Zenghui, et al.
Pubblicazione: (2025)
Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels
di: Du, Chenghao, et al.
Pubblicazione: (2025)
di: Du, Chenghao, et al.
Pubblicazione: (2025)
RobustMask: Certified Robustness against Adversarial Neural Ranking Attack via Randomized Masking
di: Liu, Jiawei, et al.
Pubblicazione: (2025)
di: Liu, Jiawei, et al.
Pubblicazione: (2025)
Certifiably Robust Image Watermark
di: Jiang, Zhengyuan, et al.
Pubblicazione: (2024)
di: Jiang, Zhengyuan, et al.
Pubblicazione: (2024)
GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis
di: Xie, Yueqi, et al.
Pubblicazione: (2024)
di: Xie, Yueqi, et al.
Pubblicazione: (2024)
Persona-Conditioned Adversarial Prompting (PCAP): Multi-Identity Red-Teaming for Enhanced Adversarial Prompt Discovery
di: Morasso, Cristian, et al.
Pubblicazione: (2026)
di: Morasso, Cristian, et al.
Pubblicazione: (2026)
Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
di: Wei, Zhang, et al.
Pubblicazione: (2025)
di: Wei, Zhang, et al.
Pubblicazione: (2025)
NCCR: to Evaluate the Robustness of Neural Networks and Adversarial Examples
di: Pu, Shi, et al.
Pubblicazione: (2025)
di: Pu, Shi, et al.
Pubblicazione: (2025)
ProxyPrompt: Securing System Prompts against Prompt Extraction Attacks
di: Zhuang, Zhixiong, et al.
Pubblicazione: (2025)
di: Zhuang, Zhixiong, et al.
Pubblicazione: (2025)
Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection
di: Chen, Meng, et al.
Pubblicazione: (2026)
di: Chen, Meng, et al.
Pubblicazione: (2026)
Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction
di: Chen, Yulin, et al.
Pubblicazione: (2025)
di: Chen, Yulin, et al.
Pubblicazione: (2025)
Robust Federated Learning Mitigates Client-side Training Data Distribution Inference Attacks
di: Xu, Yichang, et al.
Pubblicazione: (2024)
di: Xu, Yichang, et al.
Pubblicazione: (2024)
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs
di: Zizzo, Giulio, et al.
Pubblicazione: (2025)
di: Zizzo, Giulio, et al.
Pubblicazione: (2025)
WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
di: Wang, Xilong, et al.
Pubblicazione: (2026)
di: Wang, Xilong, et al.
Pubblicazione: (2026)
Documenti analoghi
-
ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data
di: Wang, Reachal, et al.
Pubblicazione: (2025) -
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
di: Zhu, Kaijie, et al.
Pubblicazione: (2025) -
SecInfer: Preventing Prompt Injection via Inference-time Scaling
di: Liu, Yupei, et al.
Pubblicazione: (2025) -
A Critical Evaluation of Defenses against Prompt Injection Attacks
di: Jia, Yuqi, et al.
Pubblicazione: (2025) -
Prompt Injection Attack to Tool Selection in LLM Agents
di: Shi, Jiawen, et al.
Pubblicazione: (2025)