Improved Generation of Adversarial Examples Against Safety-aligned LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Qizhang, Guo, Yiwen, Zuo, Wangmeng, Chen, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Transferability of Adversarial Examples via Bayesian Attacks
by: Li, Qizhang, et al.
Published: (2023)
by: Li, Qizhang, et al.
Published: (2023)
Deciphering the Chaos: Enhancing Jailbreak Attacks via Adversarial Prompt Translation
by: Li, Qizhang, et al.
Published: (2024)
by: Li, Qizhang, et al.
Published: (2024)
Position: Towards Resilience Against Adversarial Examples
by: Dai, Sihui, et al.
Published: (2024)
by: Dai, Sihui, et al.
Published: (2024)
Detecting Adversarial Examples
by: Mumcu, Furkan, et al.
Published: (2024)
by: Mumcu, Furkan, et al.
Published: (2024)
Understanding Deep Learning defenses Against Adversarial Examples Through Visualizations for Dynamic Risk Assessment
by: Echeberria-Barrio, Xabier, et al.
Published: (2024)
by: Echeberria-Barrio, Xabier, et al.
Published: (2024)
Transferability Ranking of Adversarial Examples
by: Levy, Mosh, et al.
Published: (2022)
by: Levy, Mosh, et al.
Published: (2022)
Laundering AI Authority with Adversarial Examples
by: Zhang, Jie, et al.
Published: (2026)
by: Zhang, Jie, et al.
Published: (2026)
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs
by: Zizzo, Giulio, et al.
Published: (2025)
by: Zizzo, Giulio, et al.
Published: (2025)
Comprehensive Survey on Adversarial Examples in Cybersecurity: Impacts, Challenges, and Mitigation Strategies
by: Li, Li
Published: (2024)
by: Li, Li
Published: (2024)
SoK: Analyzing Adversarial Examples: A Framework to Study Adversary Knowledge
by: Fenaux, Lucas, et al.
Published: (2024)
by: Fenaux, Lucas, et al.
Published: (2024)
AED-PADA:Improving Generalizability of Adversarial Example Detection via Principal Adversarial Domain Adaptation
by: Peng, Heqi, et al.
Published: (2024)
by: Peng, Heqi, et al.
Published: (2024)
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
Constructing Semantics-Aware Adversarial Examples with a Probabilistic Perspective
by: Zhang, Andi, et al.
Published: (2023)
by: Zhang, Andi, et al.
Published: (2023)
When and How to Fool Explainable Models (and Humans) with Adversarial Examples
by: Vadillo, Jon, et al.
Published: (2021)
by: Vadillo, Jon, et al.
Published: (2021)
Model-Editing-Based Jailbreak against Safety-aligned Large Language Models
by: Li, Yuxi, et al.
Published: (2024)
by: Li, Yuxi, et al.
Published: (2024)
TA3: Testing Against Adversarial Attacks on Machine Learning Models
by: Jin, Yuanzhe, et al.
Published: (2024)
by: Jin, Yuanzhe, et al.
Published: (2024)
Understanding and Improving Continuous Adversarial Training for LLMs via In-context Learning Theory
by: Fu, Shaopeng, et al.
Published: (2026)
by: Fu, Shaopeng, et al.
Published: (2026)
Robustness Against Adversarial Attacks via Learning Confined Adversarial Polytopes
by: Hamidi, Shayan Mohajer, et al.
Published: (2024)
by: Hamidi, Shayan Mohajer, et al.
Published: (2024)
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
by: Brown, Hannah, et al.
Published: (2024)
by: Brown, Hannah, et al.
Published: (2024)
Constructing Adversarial Examples for Vertical Federated Learning: Optimal Client Corruption through Multi-Armed Bandit
by: Yao, Duanyi, et al.
Published: (2024)
by: Yao, Duanyi, et al.
Published: (2024)
Instruction Backdoor Attacks Against Customized LLMs
by: Zhang, Rui, et al.
Published: (2024)
by: Zhang, Rui, et al.
Published: (2024)
Certifiable Black-Box Attacks with Randomized Adversarial Examples: Breaking Defenses with Provable Confidence
by: Hong, Hanbin, et al.
Published: (2023)
by: Hong, Hanbin, et al.
Published: (2023)
TaeBench: Improving Quality of Toxic Adversarial Examples
by: Zhu, Xuan, et al.
Published: (2024)
by: Zhu, Xuan, et al.
Published: (2024)
Game-Theoretic Unlearnable Example Generator
by: Liu, Shuang, et al.
Published: (2024)
by: Liu, Shuang, et al.
Published: (2024)
Game of Trojans: Adaptive Adversaries Against Output-based Trojaned-Model Detectors
by: Sahabandu, Dinuka, et al.
Published: (2024)
by: Sahabandu, Dinuka, et al.
Published: (2024)
Adversarial Sparse Teacher: Defense Against Distillation-Based Model Stealing Attacks Using Adversarial Examples
by: Yilmaz, Eda, et al.
Published: (2024)
by: Yilmaz, Eda, et al.
Published: (2024)
Comments on "Privacy-Enhanced Federated Learning Against Poisoning Adversaries"
by: Schneider, Thomas, et al.
Published: (2024)
by: Schneider, Thomas, et al.
Published: (2024)
Exploring DNN Robustness Against Adversarial Attacks Using Approximate Multipliers
by: Askarizadeh, Mohammad Javad, et al.
Published: (2024)
by: Askarizadeh, Mohammad Javad, et al.
Published: (2024)
Stealing the Invisible: Unveiling Pre-Trained CNN Models through Adversarial Examples and Timing Side-Channels
by: Shukla, Shubhi, et al.
Published: (2024)
by: Shukla, Shubhi, et al.
Published: (2024)
Efficient Adversarial Training in LLMs with Continuous Attacks
by: Xhonneux, Sophie, et al.
Published: (2024)
by: Xhonneux, Sophie, et al.
Published: (2024)
LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
by: Li, Ran, et al.
Published: (2025)
by: Li, Ran, et al.
Published: (2025)
CARE: Ensemble Adversarial Robustness Evaluation Against Adaptive Attackers for Security Applications
by: Zhang, Hangsheng, et al.
Published: (2024)
by: Zhang, Hangsheng, et al.
Published: (2024)
Adversarial Attacks Against Deep Learning-Based Radio Frequency Fingerprint Identification
by: Ma, Jie, et al.
Published: (2025)
by: Ma, Jie, et al.
Published: (2025)
PEAS: A Strategy for Crafting Transferable Adversarial Examples
by: Avraham, Bar, et al.
Published: (2024)
by: Avraham, Bar, et al.
Published: (2024)
Adversarial Suffix Filtering: a Defense Pipeline for LLMs
by: Khachaturov, David, et al.
Published: (2025)
by: Khachaturov, David, et al.
Published: (2025)
Transferable Adversarial Examples with Bayes Approach
by: Fan, Mingyuan, et al.
Published: (2022)
by: Fan, Mingyuan, et al.
Published: (2022)
Provably Unlearnable Data Examples
by: Wang, Derui, et al.
Published: (2024)
by: Wang, Derui, et al.
Published: (2024)
Rectifying Adversarial Examples Using Their Vulnerabilities
by: Morimoto, Fumiya, et al.
Published: (2026)
by: Morimoto, Fumiya, et al.
Published: (2026)
RAMP: Boosting Adversarial Robustness Against Multiple $l_p$ Perturbations for Universal Robustness
by: Jiang, Enyi, et al.
Published: (2024)
by: Jiang, Enyi, et al.
Published: (2024)
SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
by: Li, Lijun, et al.
Published: (2024)
by: Li, Lijun, et al.
Published: (2024)
Similar Items
-
Improving Transferability of Adversarial Examples via Bayesian Attacks
by: Li, Qizhang, et al.
Published: (2023) -
Deciphering the Chaos: Enhancing Jailbreak Attacks via Adversarial Prompt Translation
by: Li, Qizhang, et al.
Published: (2024) -
Position: Towards Resilience Against Adversarial Examples
by: Dai, Sihui, et al.
Published: (2024) -
Detecting Adversarial Examples
by: Mumcu, Furkan, et al.
Published: (2024) -
Understanding Deep Learning defenses Against Adversarial Examples Through Visualizations for Dynamic Risk Assessment
by: Echeberria-Barrio, Xabier, et al.
Published: (2024)