A Closer Look at Adversarial Suffix Learning for Jailbreaking LLMs: Augmented Adversarial Trigger Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zhe, Qi, Yanjun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMs
by: Basani, Advik Raj, et al.
Published: (2024)
by: Basani, Advik Raj, et al.
Published: (2024)
Adversarial Reasoning at Jailbreaking Time
by: Sabbaghi, Mahdi, et al.
Published: (2025)
by: Sabbaghi, Mahdi, et al.
Published: (2025)
A Closer Look at the Application of Causal Inference in Graph Representation Learning
by: Gao, Hang, et al.
Published: (2026)
by: Gao, Hang, et al.
Published: (2026)
Model-Based Offline Reinforcement Learning with Adversarial Data Augmentation
by: Cao, Hongye, et al.
Published: (2025)
by: Cao, Hongye, et al.
Published: (2025)
Adversarial Curriculum Graph Contrastive Learning with Pair-wise Augmentation
by: Zhao, Xinjian, et al.
Published: (2024)
by: Zhao, Xinjian, et al.
Published: (2024)
A Closer Look at Personalized Fine-Tuning in Heterogeneous Federated Learning
by: Chen, Minghui, et al.
Published: (2025)
by: Chen, Minghui, et al.
Published: (2025)
On the Implicit Adversariality of Catastrophic Forgetting in Deep Continual Learning
by: Peng, Ze, et al.
Published: (2025)
by: Peng, Ze, et al.
Published: (2025)
The Struggle Between Continuation and Refusal: A Mechanistic Analysis of the Continuation-Triggered Jailbreak in LLMs
by: Deng, Yonghong, et al.
Published: (2026)
by: Deng, Yonghong, et al.
Published: (2026)
How Does the Smoothness Approximation Method Facilitate Generalization for Federated Adversarial Learning?
by: Ding, Wenjun, et al.
Published: (2024)
by: Ding, Wenjun, et al.
Published: (2024)
Disentangling Policy from Offline Task Representation Learning via Adversarial Data Augmentation
by: Jia, Chengxing, et al.
Published: (2024)
by: Jia, Chengxing, et al.
Published: (2024)
Adversarial Preference Learning for Robust LLM Alignment
by: Wang, Yuanfu, et al.
Published: (2025)
by: Wang, Yuanfu, et al.
Published: (2025)
Diffusion-Reward Adversarial Imitation Learning
by: Lai, Chun-Mao, et al.
Published: (2024)
by: Lai, Chun-Mao, et al.
Published: (2024)
Adversarial Diffusion for Robust Reinforcement Learning
by: Foffano, Daniele, et al.
Published: (2025)
by: Foffano, Daniele, et al.
Published: (2025)
Algorithms for Adversarially Robust Deep Learning
by: Robey, Alexander
Published: (2025)
by: Robey, Alexander
Published: (2025)
Maintaining Adversarial Robustness in Continuous Learning
by: Ru, Xiaolei, et al.
Published: (2024)
by: Ru, Xiaolei, et al.
Published: (2024)
Adversarial Imitation Learning via Boosting
by: Chang, Jonathan D., et al.
Published: (2024)
by: Chang, Jonathan D., et al.
Published: (2024)
Sample-efficient Adversarial Imitation Learning
by: Jung, Dahuin, et al.
Published: (2023)
by: Jung, Dahuin, et al.
Published: (2023)
Scale-free Adversarial Reinforcement Learning
by: Chen, Mingyu, et al.
Published: (2024)
by: Chen, Mingyu, et al.
Published: (2024)
A Closer Look at Machine Unlearning for Large Language Models
by: Yuan, Xiaojian, et al.
Published: (2024)
by: Yuan, Xiaojian, et al.
Published: (2024)
Fight Back Against Jailbreaking via Prompt Adversarial Tuning
by: Mo, Yichuan, et al.
Published: (2024)
by: Mo, Yichuan, et al.
Published: (2024)
A Closer Look at Multimodal Representation Collapse
by: Chaudhuri, Abhra, et al.
Published: (2025)
by: Chaudhuri, Abhra, et al.
Published: (2025)
Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization
by: Tang, Haochun, et al.
Published: (2026)
by: Tang, Haochun, et al.
Published: (2026)
Look Closer! An Adversarial Parametric Editing Framework for Hallucination Mitigation in VLMs
by: Hu, Jiayu, et al.
Published: (2025)
by: Hu, Jiayu, et al.
Published: (2025)
Online Adversarial Knowledge Distillation for Graph Neural Networks
by: Wang, Can, et al.
Published: (2021)
by: Wang, Can, et al.
Published: (2021)
TurboFuzzLLM: Turbocharging Mutation-based Fuzzing for Effectively Jailbreaking Large Language Models in Practice
by: Goel, Aman, et al.
Published: (2025)
by: Goel, Aman, et al.
Published: (2025)
Adversarial Signed Graph Learning with Differential Privacy
by: Ke, Haobin, et al.
Published: (2025)
by: Ke, Haobin, et al.
Published: (2025)
Towards Adversarially Robust Deep Metric Learning
by: Ke, Xiaopeng
Published: (2025)
by: Ke, Xiaopeng
Published: (2025)
Regret-Based Defense in Adversarial Reinforcement Learning
by: Belaire, Roman, et al.
Published: (2023)
by: Belaire, Roman, et al.
Published: (2023)
TaeBench: Improving Quality of Toxic Adversarial Examples
by: Zhu, Xuan, et al.
Published: (2024)
by: Zhu, Xuan, et al.
Published: (2024)
Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding
by: Ryu, Hyun, et al.
Published: (2024)
by: Ryu, Hyun, et al.
Published: (2024)
GAR: Generative Adversarial Reinforcement Learning for Formal Theorem Proving
by: Wang, Ruida, et al.
Published: (2025)
by: Wang, Ruida, et al.
Published: (2025)
Adversarial Machine Learning: Bayesian Perspectives
by: Insua, David Rios, et al.
Published: (2020)
by: Insua, David Rios, et al.
Published: (2020)
Learning to Attack: A Bandit Approach to Adversarial Context Poisoning
by: Telikani, Ray, et al.
Published: (2026)
by: Telikani, Ray, et al.
Published: (2026)
Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
by: Liu, Qihao, et al.
Published: (2025)
by: Liu, Qihao, et al.
Published: (2025)
Robust Deep Reinforcement Learning with Adaptive Adversarial Perturbations in Action Space
by: Liu, Qianmei, et al.
Published: (2024)
by: Liu, Qianmei, et al.
Published: (2024)
Causal-Aware Generative Adversarial Networks with Reinforcement Learning
by: Nguyen, Tu Anh Hoang, et al.
Published: (2025)
by: Nguyen, Tu Anh Hoang, et al.
Published: (2025)
Diffusion Guided Adversarial State Perturbations in Reinforcement Learning
by: Sun, Xiaolin, et al.
Published: (2025)
by: Sun, Xiaolin, et al.
Published: (2025)
Adversarial Reinforcement Learning Framework for ESP Cheater Simulation
by: Park, Inkyu, et al.
Published: (2025)
by: Park, Inkyu, et al.
Published: (2025)
An Adversarial Learning Approach to Irregular Time-Series Forecasting
by: Nam, Heejeong, et al.
Published: (2024)
by: Nam, Heejeong, et al.
Published: (2024)
Analyzing the Impact of Adversarial Examples on Explainable Machine Learning
by: Devabhakthini, Prathyusha, et al.
Published: (2023)
by: Devabhakthini, Prathyusha, et al.
Published: (2023)
Similar Items
-
GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMs
by: Basani, Advik Raj, et al.
Published: (2024) -
Adversarial Reasoning at Jailbreaking Time
by: Sabbaghi, Mahdi, et al.
Published: (2025) -
A Closer Look at the Application of Causal Inference in Graph Representation Learning
by: Gao, Hang, et al.
Published: (2026) -
Model-Based Offline Reinforcement Learning with Adversarial Data Augmentation
by: Cao, Hongye, et al.
Published: (2025) -
Adversarial Curriculum Graph Contrastive Learning with Pair-wise Augmentation
by: Zhao, Xinjian, et al.
Published: (2024)