Persona-Conditioned Adversarial Prompting (PCAP): Multi-Identity Red-Teaming for Enhanced Adversarial Prompt Discovery
Fuente:
arXiv
Salvato in:
| Autori principali: | Morasso, Cristian, Halimi, Anisa, Hameed, Muhammad Zaid, Leith, Douglas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Persona-Conditioned Adversarial Prompting: Multi-Identity Red-Teaming for Adversarial Discovery and Mitigation
di: Morasso, Cristian, et al.
Pubblicazione: (2026)
di: Morasso, Cristian, et al.
Pubblicazione: (2026)
Towards a Re-evaluation of Data Forging Attacks in Practice
di: Suliman, Mohamed, et al.
Pubblicazione: (2024)
di: Suliman, Mohamed, et al.
Pubblicazione: (2024)
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs
di: Zizzo, Giulio, et al.
Pubblicazione: (2025)
di: Zizzo, Giulio, et al.
Pubblicazione: (2025)
Autonomous Adversary: Red-Teaming in the age of LLM
di: Mamun, Mohammad, et al.
Pubblicazione: (2026)
di: Mamun, Mohammad, et al.
Pubblicazione: (2026)
Prompt Optimization and Evaluation for LLM Automated Red Teaming
di: Freenor, Michael, et al.
Pubblicazione: (2025)
di: Freenor, Michael, et al.
Pubblicazione: (2025)
LoMime: Query-Efficient Membership Inference using Model Extraction in Label-Only Settings
di: Oksuz, Abdullah Caglar, et al.
Pubblicazione: (2026)
di: Oksuz, Abdullah Caglar, et al.
Pubblicazione: (2026)
AUTOLYCUS: Exploiting Explainable AI (XAI) for Model Extraction Attacks against Interpretable Models
di: Oksuz, Abdullah Caglar, et al.
Pubblicazione: (2023)
di: Oksuz, Abdullah Caglar, et al.
Pubblicazione: (2023)
Red Team Redemption: A Structured Comparison of Open-Source Tools for Adversary Emulation
di: Landauer, Max, et al.
Pubblicazione: (2024)
di: Landauer, Max, et al.
Pubblicazione: (2024)
Poison Attacks and Adversarial Prompts Against an Informed University Virtual Assistant
di: Fernandez, Ivan A., et al.
Pubblicazione: (2024)
di: Fernandez, Ivan A., et al.
Pubblicazione: (2024)
MoJE: Mixture of Jailbreak Experts, Naive Tabular Classifiers as Guard for Prompt Attacks
di: Cornacchia, Giandomenico, et al.
Pubblicazione: (2024)
di: Cornacchia, Giandomenico, et al.
Pubblicazione: (2024)
FlashRT: Towards Computationally and Memory Efficient Red-Teaming for Prompt Injection and Knowledge Corruption
di: Wang, Yanting, et al.
Pubblicazione: (2026)
di: Wang, Yanting, et al.
Pubblicazione: (2026)
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
di: Zhang, Zheng, et al.
Pubblicazione: (2025)
di: Zhang, Zheng, et al.
Pubblicazione: (2025)
PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses
di: Yin, Chenlong, et al.
Pubblicazione: (2026)
di: Yin, Chenlong, et al.
Pubblicazione: (2026)
Defending Jailbreak Prompts via In-Context Adversarial Game
di: Zhou, Yujun, et al.
Pubblicazione: (2024)
di: Zhou, Yujun, et al.
Pubblicazione: (2024)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
di: Reddy, Aashray, et al.
Pubblicazione: (2025)
di: Reddy, Aashray, et al.
Pubblicazione: (2025)
Enhancing Adversarial Transferability with Adversarial Weight Tuning
di: Chen, Jiahao, et al.
Pubblicazione: (2024)
di: Chen, Jiahao, et al.
Pubblicazione: (2024)
Deciphering the Chaos: Enhancing Jailbreak Attacks via Adversarial Prompt Translation
di: Li, Qizhang, et al.
Pubblicazione: (2024)
di: Li, Qizhang, et al.
Pubblicazione: (2024)
Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
di: Wei, Zhang, et al.
Pubblicazione: (2025)
di: Wei, Zhang, et al.
Pubblicazione: (2025)
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
di: Du, Pengfei
Pubblicazione: (2025)
di: Du, Pengfei
Pubblicazione: (2025)
Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection
di: Lin, Lixing, et al.
Pubblicazione: (2026)
di: Lin, Lixing, et al.
Pubblicazione: (2026)
PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
di: Zhu, Kaijie, et al.
Pubblicazione: (2023)
di: Zhu, Kaijie, et al.
Pubblicazione: (2023)
Whispers of Wealth: Red-Teaming Google's Agent Payments Protocol via Prompt Injection
di: Debi, Tanusree, et al.
Pubblicazione: (2026)
di: Debi, Tanusree, et al.
Pubblicazione: (2026)
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
di: Cao, Tri, et al.
Pubblicazione: (2026)
di: Cao, Tri, et al.
Pubblicazione: (2026)
Blockchain-Enabled Device-Enhanced Multi-Access Edge Computing in Open Adversarial Environments
di: Islam, Muhammad, et al.
Pubblicazione: (2024)
di: Islam, Muhammad, et al.
Pubblicazione: (2024)
Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels
di: Du, Chenghao, et al.
Pubblicazione: (2025)
di: Du, Chenghao, et al.
Pubblicazione: (2025)
Query-Based Adversarial Prompt Generation
di: Hayase, Jonathan, et al.
Pubblicazione: (2024)
di: Hayase, Jonathan, et al.
Pubblicazione: (2024)
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
di: Syros, Georgios, et al.
Pubblicazione: (2026)
di: Syros, Georgios, et al.
Pubblicazione: (2026)
Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
di: Pathade, Chetan
Pubblicazione: (2025)
di: Pathade, Chetan
Pubblicazione: (2025)
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
di: Maloyan, Narek, et al.
Pubblicazione: (2025)
di: Maloyan, Narek, et al.
Pubblicazione: (2025)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
di: Ma, Jiachen, et al.
Pubblicazione: (2024)
di: Ma, Jiachen, et al.
Pubblicazione: (2024)
A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection
di: Zhang, Ivan
Pubblicazione: (2025)
di: Zhang, Ivan
Pubblicazione: (2025)
LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment Training
di: Gong, Yuyang, et al.
Pubblicazione: (2026)
di: Gong, Yuyang, et al.
Pubblicazione: (2026)
How Secure is Secure Code Generation? Adversarial Prompts Put LLM Defenses to the Test
di: Tessa, Melissa, et al.
Pubblicazione: (2026)
di: Tessa, Melissa, et al.
Pubblicazione: (2026)
SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts
di: Xin, Yuan, et al.
Pubblicazione: (2026)
di: Xin, Yuan, et al.
Pubblicazione: (2026)
Certifying LLM Safety against Adversarial Prompting
di: Kumar, Aounon, et al.
Pubblicazione: (2023)
di: Kumar, Aounon, et al.
Pubblicazione: (2023)
Sentra-Guard: A Real-Time Multilingual Defense Against Adversarial LLM Prompts
di: Hasan, Md. Mehedi, et al.
Pubblicazione: (2025)
di: Hasan, Md. Mehedi, et al.
Pubblicazione: (2025)
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
di: Li, Xiang, et al.
Pubblicazione: (2025)
di: Li, Xiang, et al.
Pubblicazione: (2025)
Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation
di: Roh, Jaechul, et al.
Pubblicazione: (2025)
di: Roh, Jaechul, et al.
Pubblicazione: (2025)
IPI-proxy: An Intercepting Proxy for Red-Teaming Web-Browsing AI Agents Against Indirect Prompt Injection
di: Chia-Pei, et al.
Pubblicazione: (2026)
di: Chia-Pei, et al.
Pubblicazione: (2026)
Improving Privacy Benefits of Redaction
di: Gusain, Vaibhav, et al.
Pubblicazione: (2025)
di: Gusain, Vaibhav, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Persona-Conditioned Adversarial Prompting: Multi-Identity Red-Teaming for Adversarial Discovery and Mitigation
di: Morasso, Cristian, et al.
Pubblicazione: (2026) -
Towards a Re-evaluation of Data Forging Attacks in Practice
di: Suliman, Mohamed, et al.
Pubblicazione: (2024) -
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs
di: Zizzo, Giulio, et al.
Pubblicazione: (2025) -
Autonomous Adversary: Red-Teaming in the age of LLM
di: Mamun, Mohammad, et al.
Pubblicazione: (2026) -
Prompt Optimization and Evaluation for LLM Automated Red Teaming
di: Freenor, Michael, et al.
Pubblicazione: (2025)