Query-Based Adversarial Prompt Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hayase, Jonathan, Borevkovic, Ema, Carlini, Nicholas, Tramèr, Florian, Nasr, Milad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025)
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025)
Are aligned neural networks adversarially aligned?
von: Carlini, Nicholas, et al.
Veröffentlicht: (2023)
von: Carlini, Nicholas, et al.
Veröffentlicht: (2023)
Universal Jailbreak Backdoors from Poisoned Human Feedback
von: Rando, Javier, et al.
Veröffentlicht: (2023)
von: Rando, Javier, et al.
Veröffentlicht: (2023)
An Adversarial Perspective on Machine Unlearning for AI Safety
von: Łucki, Jakub, et al.
Veröffentlicht: (2024)
von: Łucki, Jakub, et al.
Veröffentlicht: (2024)
Stealing User Prompts from Mixture of Experts
von: Yona, Itay, et al.
Veröffentlicht: (2024)
von: Yona, Itay, et al.
Veröffentlicht: (2024)
Remote Timing Attacks on Efficient Language Model Inference
von: Carlini, Nicholas, et al.
Veröffentlicht: (2024)
von: Carlini, Nicholas, et al.
Veröffentlicht: (2024)
SoK: Watermarking for AI-Generated Content
von: Zhao, Xuandong, et al.
Veröffentlicht: (2024)
von: Zhao, Xuandong, et al.
Veröffentlicht: (2024)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)
Large-scale online deanonymization with LLMs
von: Lermen, Simon, et al.
Veröffentlicht: (2026)
von: Lermen, Simon, et al.
Veröffentlicht: (2026)
Adversarial ML Problems Are Getting Harder to Solve and to Evaluate
von: Rando, Javier, et al.
Veröffentlicht: (2025)
von: Rando, Javier, et al.
Veröffentlicht: (2025)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
von: Rando, Javier, et al.
Veröffentlicht: (2024)
von: Rando, Javier, et al.
Veröffentlicht: (2024)
Certifying LLM Safety against Adversarial Prompting
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
von: Paulus, Anselm, et al.
Veröffentlicht: (2024)
von: Paulus, Anselm, et al.
Veröffentlicht: (2024)
Fight Back Against Jailbreaking via Prompt Adversarial Tuning
von: Mo, Yichuan, et al.
Veröffentlicht: (2024)
von: Mo, Yichuan, et al.
Veröffentlicht: (2024)
TRAP: Targeted Random Adversarial Prompt Honeypot for Black-Box Identification
von: Gubri, Martin, et al.
Veröffentlicht: (2024)
von: Gubri, Martin, et al.
Veröffentlicht: (2024)
Evading Black-box Classifiers Without Breaking Eggs
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2023)
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2023)
Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining
von: Tramèr, Florian, et al.
Veröffentlicht: (2022)
von: Tramèr, Florian, et al.
Veröffentlicht: (2022)
RECAP: A Resource-Efficient Method for Adversarial Prompting in Large Language Models
von: Chugh, Rishit
Veröffentlicht: (2026)
von: Chugh, Rishit
Veröffentlicht: (2026)
LLMs unlock new paths to monetizing exploits
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025)
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025)
Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting
von: To, Bang Trinh Tran, et al.
Veröffentlicht: (2025)
von: To, Bang Trinh Tran, et al.
Veröffentlicht: (2025)
Adversarial Attacks on Parts of Speech: An Empirical Study in Text-to-Image Generation
von: Shahariar, G M, et al.
Veröffentlicht: (2024)
von: Shahariar, G M, et al.
Veröffentlicht: (2024)
On Evaluating The Performance of Watermarked Machine-Generated Texts Under Adversarial Attacks
von: Liu, Zesen, et al.
Veröffentlicht: (2024)
von: Liu, Zesen, et al.
Veröffentlicht: (2024)
IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
von: Guo, Chuan, et al.
Veröffentlicht: (2026)
von: Guo, Chuan, et al.
Veröffentlicht: (2026)
Blind Baselines Beat Membership Inference Attacks for Foundation Models
von: Das, Debeshee, et al.
Veröffentlicht: (2024)
von: Das, Debeshee, et al.
Veröffentlicht: (2024)
Privacy Side Channels in Machine Learning Systems
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2023)
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2023)
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
von: Liang, Buyun, et al.
Veröffentlicht: (2025)
von: Liang, Buyun, et al.
Veröffentlicht: (2025)
Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
von: Yang, Xiaoxue, et al.
Veröffentlicht: (2025)
von: Yang, Xiaoxue, et al.
Veröffentlicht: (2025)
Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards
von: Huang, Yangsibo, et al.
Veröffentlicht: (2025)
von: Huang, Yangsibo, et al.
Veröffentlicht: (2025)
Prompt2Fingerprint: Plug-and-Play LLM Fingerprinting via Text-to-Weight Generation
von: Chen, Sixu, et al.
Veröffentlicht: (2026)
von: Chen, Sixu, et al.
Veröffentlicht: (2026)
PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
von: Geng, Runpeng, et al.
Veröffentlicht: (2025)
von: Geng, Runpeng, et al.
Veröffentlicht: (2025)
Jailbreaking with Universal Multi-Prompts
von: Hsu, Yu-Ling, et al.
Veröffentlicht: (2025)
von: Hsu, Yu-Ling, et al.
Veröffentlicht: (2025)
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
von: Nikolić, Kristina, et al.
Veröffentlicht: (2025)
von: Nikolić, Kristina, et al.
Veröffentlicht: (2025)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
von: Saiem, Bijoy Ahmed, et al.
Veröffentlicht: (2024)
von: Saiem, Bijoy Ahmed, et al.
Veröffentlicht: (2024)
On Adversarial Robustness of Language Models in Transfer Learning
von: Turbal, Bohdan, et al.
Veröffentlicht: (2024)
von: Turbal, Bohdan, et al.
Veröffentlicht: (2024)
TaeBench: Improving Quality of Toxic Adversarial Examples
von: Zhu, Xuan, et al.
Veröffentlicht: (2024)
von: Zhu, Xuan, et al.
Veröffentlicht: (2024)
The Resurgence of GCG Adversarial Attacks on Large Language Models
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
PIArena: A Platform for Prompt Injection Evaluation
von: Geng, Runpeng, et al.
Veröffentlicht: (2026)
von: Geng, Runpeng, et al.
Veröffentlicht: (2026)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
von: Liu, Yupei, et al.
Veröffentlicht: (2023)
von: Liu, Yupei, et al.
Veröffentlicht: (2023)
Adversarial Vulnerabilities in Large Language Models for Time Series Forecasting
von: Liu, Fuqiang, et al.
Veröffentlicht: (2024)
von: Liu, Fuqiang, et al.
Veröffentlicht: (2024)
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
von: Liang, Buyun, et al.
Veröffentlicht: (2026)
von: Liang, Buyun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025) -
Are aligned neural networks adversarially aligned?
von: Carlini, Nicholas, et al.
Veröffentlicht: (2023) -
Universal Jailbreak Backdoors from Poisoned Human Feedback
von: Rando, Javier, et al.
Veröffentlicht: (2023) -
An Adversarial Perspective on Machine Unlearning for AI Safety
von: Łucki, Jakub, et al.
Veröffentlicht: (2024) -
Stealing User Prompts from Mixture of Experts
von: Yona, Itay, et al.
Veröffentlicht: (2024)