Persona-Conditioned Adversarial Prompting: Multi-Identity Red-Teaming for Adversarial Discovery and Mitigation
Fuente:
arXiv
Guardado en:
| Autores principales: | Morasso, Cristian, Halimi, Anisa, Hameed, Muhammad Zaid, Leith, Douglas |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Persona-Conditioned Adversarial Prompting (PCAP): Multi-Identity Red-Teaming for Enhanced Adversarial Prompt Discovery
por: Morasso, Cristian, et al.
Publicado: (2026)
por: Morasso, Cristian, et al.
Publicado: (2026)
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs
por: Zizzo, Giulio, et al.
Publicado: (2025)
por: Zizzo, Giulio, et al.
Publicado: (2025)
LoMime: Query-Efficient Membership Inference using Model Extraction in Label-Only Settings
por: Oksuz, Abdullah Caglar, et al.
Publicado: (2026)
por: Oksuz, Abdullah Caglar, et al.
Publicado: (2026)
AUTOLYCUS: Exploiting Explainable AI (XAI) for Model Extraction Attacks against Interpretable Models
por: Oksuz, Abdullah Caglar, et al.
Publicado: (2023)
por: Oksuz, Abdullah Caglar, et al.
Publicado: (2023)
Towards a Re-evaluation of Data Forging Attacks in Practice
por: Suliman, Mohamed, et al.
Publicado: (2024)
por: Suliman, Mohamed, et al.
Publicado: (2024)
Mitigating Error Amplification in Fast Adversarial Training
por: Zhao, Mengnan, et al.
Publicado: (2026)
por: Zhao, Mengnan, et al.
Publicado: (2026)
Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM Agents
por: He, Pengfei, et al.
Publicado: (2026)
por: He, Pengfei, et al.
Publicado: (2026)
Improving Privacy Benefits of Redaction
por: Gusain, Vaibhav, et al.
Publicado: (2025)
por: Gusain, Vaibhav, et al.
Publicado: (2025)
Attack Atlas: A Practitioner's Perspective on Challenges and Pitfalls in Red Teaming GenAI
por: Rawat, Ambrish, et al.
Publicado: (2024)
por: Rawat, Ambrish, et al.
Publicado: (2024)
PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses
por: Yin, Chenlong, et al.
Publicado: (2026)
por: Yin, Chenlong, et al.
Publicado: (2026)
MoJE: Mixture of Jailbreak Experts, Naive Tabular Classifiers as Guard for Prompt Attacks
por: Cornacchia, Giandomenico, et al.
Publicado: (2024)
por: Cornacchia, Giandomenico, et al.
Publicado: (2024)
Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards
por: Huang, Yangsibo, et al.
Publicado: (2025)
por: Huang, Yangsibo, et al.
Publicado: (2025)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
por: Reddy, Aashray, et al.
Publicado: (2025)
por: Reddy, Aashray, et al.
Publicado: (2025)
Defending Jailbreak Prompts via In-Context Adversarial Game
por: Zhou, Yujun, et al.
Publicado: (2024)
por: Zhou, Yujun, et al.
Publicado: (2024)
Fine-Tuning Personalization in Federated Learning to Mitigate Adversarial Clients
por: Allouah, Youssef, et al.
Publicado: (2024)
por: Allouah, Youssef, et al.
Publicado: (2024)
On the Robustness of Malware Detectors to Adversarial Samples
por: Salman, Muhammad, et al.
Publicado: (2024)
por: Salman, Muhammad, et al.
Publicado: (2024)
Comprehensive Survey on Adversarial Examples in Cybersecurity: Impacts, Challenges, and Mitigation Strategies
por: Li, Li
Publicado: (2024)
por: Li, Li
Publicado: (2024)
Adversarial Sample Generation for Anomaly Detection in Industrial Control Systems
por: Mustafa, Abdul, et al.
Publicado: (2025)
por: Mustafa, Abdul, et al.
Publicado: (2025)
Mitigating the Structural Bias in Graph Adversarial Defenses
por: Fang, Junyuan, et al.
Publicado: (2025)
por: Fang, Junyuan, et al.
Publicado: (2025)
Detecting Adversarial Data via Provable Adversarial Noise Amplification
por: Mumcu, Furkan, et al.
Publicado: (2026)
por: Mumcu, Furkan, et al.
Publicado: (2026)
Enhancing Adversarial Attacks via Parameter Adaptive Adversarial Attack
por: Jin, Zhibo, et al.
Publicado: (2024)
por: Jin, Zhibo, et al.
Publicado: (2024)
Stealthy Adversarial Attacks on Stochastic Multi-Armed Bandits
por: Wang, Zhiwei, et al.
Publicado: (2024)
por: Wang, Zhiwei, et al.
Publicado: (2024)
Double-Adversarial Activation Anomaly Detection: Adversarial Autoencoders are Anomaly Generators
por: Schulze, J. -P., et al.
Publicado: (2021)
por: Schulze, J. -P., et al.
Publicado: (2021)
Robustness Against Adversarial Attacks via Learning Confined Adversarial Polytopes
por: Hamidi, Shayan Mohajer, et al.
Publicado: (2024)
por: Hamidi, Shayan Mohajer, et al.
Publicado: (2024)
Generalization Properties of Adversarial Training for $\ell_0$-Bounded Adversarial Attacks
por: Delgosha, Payam, et al.
Publicado: (2024)
por: Delgosha, Payam, et al.
Publicado: (2024)
Detecting Adversarial Examples
por: Mumcu, Furkan, et al.
Publicado: (2024)
por: Mumcu, Furkan, et al.
Publicado: (2024)
Adversarial Machine Unlearning
por: Di, Zonglin, et al.
Publicado: (2024)
por: Di, Zonglin, et al.
Publicado: (2024)
SoK: Analyzing Adversarial Examples: A Framework to Study Adversary Knowledge
por: Fenaux, Lucas, et al.
Publicado: (2024)
por: Fenaux, Lucas, et al.
Publicado: (2024)
Autonomous Adversary: Red-Teaming in the age of LLM
por: Mamun, Mohammad, et al.
Publicado: (2026)
por: Mamun, Mohammad, et al.
Publicado: (2026)
Test-time Adversarial Defense with Opposite Adversarial Path and High Attack Time Cost
por: Yeh, Cheng-Han, et al.
Publicado: (2024)
por: Yeh, Cheng-Han, et al.
Publicado: (2024)
PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
por: Zhu, Kaijie, et al.
Publicado: (2023)
por: Zhu, Kaijie, et al.
Publicado: (2023)
Adaptive Randomized Smoothing: Certified Adversarial Robustness for Multi-Step Defences
por: Lyu, Saiyue, et al.
Publicado: (2024)
por: Lyu, Saiyue, et al.
Publicado: (2024)
Adversarial Observations in Weather Forecasting
por: Imgrund, Erik, et al.
Publicado: (2025)
por: Imgrund, Erik, et al.
Publicado: (2025)
Transferability Ranking of Adversarial Examples
por: Levy, Mosh, et al.
Publicado: (2022)
por: Levy, Mosh, et al.
Publicado: (2022)
AMUN: Adversarial Machine UNlearning
por: Ebrahimpour-Boroojeny, Ali, et al.
Publicado: (2025)
por: Ebrahimpour-Boroojeny, Ali, et al.
Publicado: (2025)
Colliding with Adversaries at ECML-PKDD 2025 Adversarial Attack Competition 1st Prize Solution
por: Stefanopoulos, Dimitris, et al.
Publicado: (2025)
por: Stefanopoulos, Dimitris, et al.
Publicado: (2025)
Poisoning the Watchtower: Prompt Injection Attacks Against LLM-Augmented Security Operations Through Adversarial Log Content
por: Pandey, Rohan, et al.
Publicado: (2026)
por: Pandey, Rohan, et al.
Publicado: (2026)
Consistent Valid Physically-Realizable Adversarial Attack against Crowd-flow Prediction Models
por: Ali, Hassan, et al.
Publicado: (2023)
por: Ali, Hassan, et al.
Publicado: (2023)
Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak Prompts
por: Liu, Yi, et al.
Publicado: (2024)
por: Liu, Yi, et al.
Publicado: (2024)
Early Approaches to Adversarial Fine-Tuning for Prompt Injection Defense: A 2022 Study of GPT-3 and Contemporary Models
por: Sandoval, Gustavo, et al.
Publicado: (2025)
por: Sandoval, Gustavo, et al.
Publicado: (2025)
Ejemplares similares
-
Persona-Conditioned Adversarial Prompting (PCAP): Multi-Identity Red-Teaming for Enhanced Adversarial Prompt Discovery
por: Morasso, Cristian, et al.
Publicado: (2026) -
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs
por: Zizzo, Giulio, et al.
Publicado: (2025) -
LoMime: Query-Efficient Membership Inference using Model Extraction in Label-Only Settings
por: Oksuz, Abdullah Caglar, et al.
Publicado: (2026) -
AUTOLYCUS: Exploiting Explainable AI (XAI) for Model Extraction Attacks against Interpretable Models
por: Oksuz, Abdullah Caglar, et al.
Publicado: (2023) -
Towards a Re-evaluation of Data Forging Attacks in Practice
por: Suliman, Mohamed, et al.
Publicado: (2024)