GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Basani, Advik Raj, Zhang, Xiao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Crafting Adversarial Inputs for Large Vision-Language Models Using Black-Box Optimization
by: Guan, Jiwei, et al.
Published: (2026)
by: Guan, Jiwei, et al.
Published: (2026)
Efficient Black-box Adversarial Attacks via Bayesian Optimization Guided by a Function Prior
by: Cheng, Shuyu, et al.
Published: (2024)
by: Cheng, Shuyu, et al.
Published: (2024)
SoK: Pitfalls in Evaluating Black-Box Attacks
by: Suya, Fnu, et al.
Published: (2023)
by: Suya, Fnu, et al.
Published: (2023)
Towards Black-Box Membership Inference Attack for Diffusion Models
by: Li, Jingwei, et al.
Published: (2024)
by: Li, Jingwei, et al.
Published: (2024)
Rewriting the Budget: A General Framework for Black-Box Attacks Under Cost Asymmetry
by: Salmani, Mahdi, et al.
Published: (2025)
by: Salmani, Mahdi, et al.
Published: (2025)
Efficient Semi-Supervised Adversarial Training via Latent Clustering-Based Data Reduction
by: Ghosh, Somrita, et al.
Published: (2025)
by: Ghosh, Somrita, et al.
Published: (2025)
FC-Attack: Jailbreaking Multimodal Large Language Models via Auto-Generated Flowcharts
by: Zhang, Ziyi, et al.
Published: (2025)
by: Zhang, Ziyi, et al.
Published: (2025)
From Attack to Defense: Insights into Deep Learning Security Measures in Black-Box Settings
by: Juraev, Firuz, et al.
Published: (2024)
by: Juraev, Firuz, et al.
Published: (2024)
Transferable Black-Box One-Shot Forging of Watermarks via Image Preference Models
by: Souček, Tomáš, et al.
Published: (2025)
by: Souček, Tomáš, et al.
Published: (2025)
Multimodal Pragmatic Jailbreak on Text-to-image Models
by: Liu, Tong, et al.
Published: (2024)
by: Liu, Tong, et al.
Published: (2024)
SemiAdv: Query-Efficient Black-Box Adversarial Attack with Unlabeled Images
by: Fan, Mingyuan, et al.
Published: (2024)
by: Fan, Mingyuan, et al.
Published: (2024)
Fingerprinting Image-to-Image Generative Adversarial Networks
by: Li, Guanlin, et al.
Published: (2021)
by: Li, Guanlin, et al.
Published: (2021)
A White-Box False Positive Adversarial Attack Method on Contrastive Loss Based Offline Handwritten Signature Verification Models
by: Guo, Zhongliang, et al.
Published: (2023)
by: Guo, Zhongliang, et al.
Published: (2023)
PuriDefense: Randomized Local Implicit Adversarial Purification for Defending Black-box Query-based Attacks
by: Guo, Ping, et al.
Published: (2024)
by: Guo, Ping, et al.
Published: (2024)
Reinforcement Learning Platform for Adversarial Black-box Attacks with Custom Distortion Filters
by: Sarkar, Soumyendu, et al.
Published: (2025)
by: Sarkar, Soumyendu, et al.
Published: (2025)
On the Importance of Backbone to the Adversarial Robustness of Object Detectors
by: Li, Xiao, et al.
Published: (2023)
by: Li, Xiao, et al.
Published: (2023)
ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test
by: Yang, Guan-Yan, et al.
Published: (2025)
by: Yang, Guan-Yan, et al.
Published: (2025)
GLEAN: Generative Learning for Eliminating Adversarial Noise
by: Kim, Justin Lyu, et al.
Published: (2024)
by: Kim, Justin Lyu, et al.
Published: (2024)
Exploring the Adversarial Frontier: Quantifying Robustness via Adversarial Hypervolume
by: Guo, Ping, et al.
Published: (2024)
by: Guo, Ping, et al.
Published: (2024)
AR-GAN: Generative Adversarial Network-Based Defense Method Against Adversarial Attacks on the Traffic Sign Classification System of Autonomous Vehicles
by: Salek, M Sabbir, et al.
Published: (2023)
by: Salek, M Sabbir, et al.
Published: (2023)
Towards Understanding Dual BN In Hybrid Adversarial Training
by: Zhang, Chenshuang, et al.
Published: (2024)
by: Zhang, Chenshuang, et al.
Published: (2024)
Data-free Defense of Black Box Models Against Adversarial Attacks
by: Nayak, Gaurav Kumar, et al.
Published: (2022)
by: Nayak, Gaurav Kumar, et al.
Published: (2022)
A Survey on the Application of Generative Adversarial Networks in Cybersecurity: Prospective, Direction and Open Research Scopes
by: Arifin, Md Mashrur, et al.
Published: (2024)
by: Arifin, Md Mashrur, et al.
Published: (2024)
X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Multi-Turn Jailbreaks without Compromising Usability
by: Lu, Xiaoya, et al.
Published: (2025)
by: Lu, Xiaoya, et al.
Published: (2025)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
by: Mehrotra, Anay, et al.
Published: (2023)
by: Mehrotra, Anay, et al.
Published: (2023)
Stop Reasoning! When Multimodal LLM with Chain-of-Thought Reasoning Meets Adversarial Image
by: Wang, Zefeng, et al.
Published: (2024)
by: Wang, Zefeng, et al.
Published: (2024)
Adversarial Detection by Approximation of Ensemble Boundary
by: Windeatt, T.
Published: (2022)
by: Windeatt, T.
Published: (2022)
DarkLLM: Learning Language-Driven Adversarial Attacks with Large Language Models
by: Sun, Ye, et al.
Published: (2026)
by: Sun, Ye, et al.
Published: (2026)
MOS-Attack: A Scalable Multi-objective Adversarial Attack Framework
by: Guo, Ping, et al.
Published: (2025)
by: Guo, Ping, et al.
Published: (2025)
BB-Patch: BlackBox Adversarial Patch-Attack using Zeroth-Order Optimization
by: Kumar, Satyadwyoom, et al.
Published: (2024)
by: Kumar, Satyadwyoom, et al.
Published: (2024)
GreedyPixel: Fine-Grained Black-Box Adversarial Attack Via Greedy Algorithm
by: Wang, Hanrui, et al.
Published: (2025)
by: Wang, Hanrui, et al.
Published: (2025)
Evaluating the Evaluators: Trust in Adversarial Robustness Tests
by: Cinà, Antonio Emanuele, et al.
Published: (2025)
by: Cinà, Antonio Emanuele, et al.
Published: (2025)
Improving the Transferability of Adversarial Attacks by an Input Transpose
by: Wan, Qing, et al.
Published: (2025)
by: Wan, Qing, et al.
Published: (2025)
The Impact of Scaling Training Data on Adversarial Robustness
by: Zimmerli, Marco, et al.
Published: (2025)
by: Zimmerli, Marco, et al.
Published: (2025)
L-AutoDA: Leveraging Large Language Models for Automated Decision-based Adversarial Attacks
by: Guo, Ping, et al.
Published: (2024)
by: Guo, Ping, et al.
Published: (2024)
Standard-Deviation-Inspired Regularization for Improving Adversarial Robustness
by: Fakorede, Olukorede, et al.
Published: (2024)
by: Fakorede, Olukorede, et al.
Published: (2024)
Impact of Architectural Modifications on Deep Learning Adversarial Robustness
by: Juraev, Firuz, et al.
Published: (2024)
by: Juraev, Firuz, et al.
Published: (2024)
Sy-FAR: Symmetry-based Fair Adversarial Robustness
by: Najjar, Haneen, et al.
Published: (2025)
by: Najjar, Haneen, et al.
Published: (2025)
Interpretation of Neural Networks is Susceptible to Universal Adversarial Perturbations
by: Oskouie, Haniyeh Ehsani, et al.
Published: (2022)
by: Oskouie, Haniyeh Ehsani, et al.
Published: (2022)
Undermining Image and Text Classification Algorithms Using Adversarial Attacks
by: Lunga, Langalibalele, et al.
Published: (2024)
by: Lunga, Langalibalele, et al.
Published: (2024)
Similar Items
-
Crafting Adversarial Inputs for Large Vision-Language Models Using Black-Box Optimization
by: Guan, Jiwei, et al.
Published: (2026) -
Efficient Black-box Adversarial Attacks via Bayesian Optimization Guided by a Function Prior
by: Cheng, Shuyu, et al.
Published: (2024) -
SoK: Pitfalls in Evaluating Black-Box Attacks
by: Suya, Fnu, et al.
Published: (2023) -
Towards Black-Box Membership Inference Attack for Diffusion Models
by: Li, Jingwei, et al.
Published: (2024) -
Rewriting the Budget: A General Framework for Black-Box Attacks Under Cost Asymmetry
by: Salmani, Mahdi, et al.
Published: (2025)