TRAP: Targeted Random Adversarial Prompt Honeypot for Black-Box Identification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gubri, Martin, Ulmer, Dennis, Lee, Hwaran, Yun, Sangdoo, Oh, Seong Joon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
von: Green, Tommaso, et al.
Veröffentlicht: (2025)
von: Green, Tommaso, et al.
Veröffentlicht: (2025)
Calibrating Large Language Models Using Their Generations Only
von: Ulmer, Dennis, et al.
Veröffentlicht: (2024)
von: Ulmer, Dennis, et al.
Veröffentlicht: (2024)
LLM in the Shell: Generative Honeypots
von: Sladić, Muris, et al.
Veröffentlicht: (2023)
von: Sladić, Muris, et al.
Veröffentlicht: (2023)
Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models
von: Puerto, Haritz, et al.
Veröffentlicht: (2024)
von: Puerto, Haritz, et al.
Veröffentlicht: (2024)
"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks
von: Wang, Libo
Veröffentlicht: (2024)
von: Wang, Libo
Veröffentlicht: (2024)
MemPot: Defending Against Memory Extraction Attack with Optimized Honeypots
von: Wang, Yuhao, et al.
Veröffentlicht: (2026)
von: Wang, Yuhao, et al.
Veröffentlicht: (2026)
C-SEO Bench: Does Conversational SEO Work?
von: Puerto, Haritz, et al.
Veröffentlicht: (2025)
von: Puerto, Haritz, et al.
Veröffentlicht: (2025)
Dr.LLM: Dynamic Layer Routing in LLMs
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
von: Akbar-Tajari, Mohammad, et al.
Veröffentlicht: (2025)
von: Akbar-Tajari, Mohammad, et al.
Veröffentlicht: (2025)
Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks
von: Gohil, Vasudev
Veröffentlicht: (2025)
von: Gohil, Vasudev
Veröffentlicht: (2025)
Adversarial Agents: Black-Box Evasion Attacks with Reinforcement Learning
von: Domico, Kyle, et al.
Veröffentlicht: (2025)
von: Domico, Kyle, et al.
Veröffentlicht: (2025)
Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test
von: Zhu, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Zhu, Xiaoyuan, et al.
Veröffentlicht: (2025)
TrailBlazer: History-Guided Reinforcement Learning for Black-Box LLM Jailbreaking
von: Yoon, Sung-Hoon, et al.
Veröffentlicht: (2026)
von: Yoon, Sung-Hoon, et al.
Veröffentlicht: (2026)
SGuard-v1: Safety Guardrail for Large Language Models
von: Lee, JoonHo, et al.
Veröffentlicht: (2025)
von: Lee, JoonHo, et al.
Veröffentlicht: (2025)
Black-Box Opinion Manipulation Attacks to Retrieval-Augmented Generation of Large Language Models
von: Chen, Zhuo, et al.
Veröffentlicht: (2024)
von: Chen, Zhuo, et al.
Veröffentlicht: (2024)
Vulnerability Disclosure through Adaptive Black-Box Adversarial Attacks on NIDS
von: Ennaji, Sabrine, et al.
Veröffentlicht: (2025)
von: Ennaji, Sabrine, et al.
Veröffentlicht: (2025)
Are All Prompt Components Value-Neutral? Understanding the Heterogeneous Adversarial Robustness of Dissected Prompt in Large Language Models
von: Zheng, Yujia, et al.
Veröffentlicht: (2025)
von: Zheng, Yujia, et al.
Veröffentlicht: (2025)
Q-FAKER: Query-free Hard Black-box Attack via Controlled Generation
von: Na, CheolWon, et al.
Veröffentlicht: (2025)
von: Na, CheolWon, et al.
Veröffentlicht: (2025)
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs
von: Berezin, Sergey, et al.
Veröffentlicht: (2025)
von: Berezin, Sergey, et al.
Veröffentlicht: (2025)
Your Inference Request Will Become a Black Box: Confidential Inference for Cloud-based Large Language Models
von: Huang, Chung-ju, et al.
Veröffentlicht: (2026)
von: Huang, Chung-ju, et al.
Veröffentlicht: (2026)
Finding a Wolf in Sheep's Clothing: Combating Adversarial Text-To-Image Prompts with Text Summarization
von: Cooper, Portia, et al.
Veröffentlicht: (2024)
von: Cooper, Portia, et al.
Veröffentlicht: (2024)
How stealthy is stealthy? Studying the Efficacy of Black-Box Adversarial Attacks in the Real World
von: Panebianco, Francesco, et al.
Veröffentlicht: (2025)
von: Panebianco, Francesco, et al.
Veröffentlicht: (2025)
Design and Development of an Intelligent LLM-based LDAP Honeypot
von: Jiménez-Román, Javier, et al.
Veröffentlicht: (2025)
von: Jiménez-Román, Javier, et al.
Veröffentlicht: (2025)
Query-Based Adversarial Prompt Generation
von: Hayase, Jonathan, et al.
Veröffentlicht: (2024)
von: Hayase, Jonathan, et al.
Veröffentlicht: (2024)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
von: Mehrotra, Anay, et al.
Veröffentlicht: (2023)
von: Mehrotra, Anay, et al.
Veröffentlicht: (2023)
Behavior-Aware and Generalizable Defense Against Black-Box Adversarial Attacks for ML-Based IDS
von: Ennaji, Sabrine, et al.
Veröffentlicht: (2025)
von: Ennaji, Sabrine, et al.
Veröffentlicht: (2025)
LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild
von: Reworr, et al.
Veröffentlicht: (2024)
von: Reworr, et al.
Veröffentlicht: (2024)
Certifying LLM Safety against Adversarial Prompting
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting
von: To, Bang Trinh Tran, et al.
Veröffentlicht: (2025)
von: To, Bang Trinh Tran, et al.
Veröffentlicht: (2025)
ArmSSL: Adversarial Robust Black-Box Watermarking for Self-Supervised Learning Pre-trained Encoders
von: Jiang, Yongqi, et al.
Veröffentlicht: (2026)
von: Jiang, Yongqi, et al.
Veröffentlicht: (2026)
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
von: Wang, Zhun, et al.
Veröffentlicht: (2025)
von: Wang, Zhun, et al.
Veröffentlicht: (2025)
PAL: Proxy-Guided Black-Box Attack on Large Language Models
von: Sitawarin, Chawin, et al.
Veröffentlicht: (2024)
von: Sitawarin, Chawin, et al.
Veröffentlicht: (2024)
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
von: Paulus, Anselm, et al.
Veröffentlicht: (2024)
von: Paulus, Anselm, et al.
Veröffentlicht: (2024)
Fight Back Against Jailbreaking via Prompt Adversarial Tuning
von: Mo, Yichuan, et al.
Veröffentlicht: (2024)
von: Mo, Yichuan, et al.
Veröffentlicht: (2024)
Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks
von: Wu, ChenYu, et al.
Veröffentlicht: (2025)
von: Wu, ChenYu, et al.
Veröffentlicht: (2025)
Comparative Analysis of Black-Box and White-Box Machine Learning Model in Phishing Detection
von: Fajar, Abdullah, et al.
Veröffentlicht: (2024)
von: Fajar, Abdullah, et al.
Veröffentlicht: (2024)
RECAP: A Resource-Efficient Method for Adversarial Prompting in Large Language Models
von: Chugh, Rishit
Veröffentlicht: (2026)
von: Chugh, Rishit
Veröffentlicht: (2026)
DePrompt: Desensitization and Evaluation of Personal Identifiable Information in Large Language Model Prompts
von: Sun, Xiongtao, et al.
Veröffentlicht: (2024)
von: Sun, Xiongtao, et al.
Veröffentlicht: (2024)
Counter-Samples: A Stateless Strategy to Neutralize Black Box Adversarial Attacks
von: Bokobza, Roey, et al.
Veröffentlicht: (2024)
von: Bokobza, Roey, et al.
Veröffentlicht: (2024)
Prompt Injection as Role Confusion
von: Ye, Charles, et al.
Veröffentlicht: (2026)
von: Ye, Charles, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
von: Green, Tommaso, et al.
Veröffentlicht: (2025) -
Calibrating Large Language Models Using Their Generations Only
von: Ulmer, Dennis, et al.
Veröffentlicht: (2024) -
LLM in the Shell: Generative Honeypots
von: Sladić, Muris, et al.
Veröffentlicht: (2023) -
Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models
von: Puerto, Haritz, et al.
Veröffentlicht: (2024) -
"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks
von: Wang, Libo
Veröffentlicht: (2024)