Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Hua, Jiaqi, Wei, Wanxu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs
by: Yoon, Sangyeon, et al.
Published: (2026)
by: Yoon, Sangyeon, et al.
Published: (2026)
SafeBehavior: Simulating Human-Like Multistage Reasoning to Mitigate Jailbreak Attacks in Large Language Models
by: Zhao, Qinjian, et al.
Published: (2025)
by: Zhao, Qinjian, et al.
Published: (2025)
Involuntary Jailbreak: On Self-Prompting Attacks
by: Guo, Yangyang, et al.
Published: (2025)
by: Guo, Yangyang, et al.
Published: (2025)
Modality-Collaborative Low-Rank Decomposers for Few-Shot Video Domain Adaptation
by: Wanyan, Yuyang, et al.
Published: (2025)
by: Wanyan, Yuyang, et al.
Published: (2025)
Enhancing Heterogeneous Knowledge Graph Completion with a Novel GAT-based Approach
by: Wei, Wanxu, et al.
Published: (2024)
by: Wei, Wanxu, et al.
Published: (2024)
Merging Improves Self-Critique Against Jailbreak Attacks
by: Gallego, Victor
Published: (2024)
by: Gallego, Victor
Published: (2024)
Advancing Graph Few-Shot Learning via In-Context Learning
by: Guan, Renchu, et al.
Published: (2026)
by: Guan, Renchu, et al.
Published: (2026)
Self-supervised Learning for Acoustic Few-Shot Classification
by: Liang, Jingyong, et al.
Published: (2024)
by: Liang, Jingyong, et al.
Published: (2024)
Learning to Detect Unseen Jailbreak Attacks in Large Vision-Language Models
by: Liang, Shuang, et al.
Published: (2025)
by: Liang, Shuang, et al.
Published: (2025)
Untargeted Jailbreak Attack
by: Huang, Xinzhe, et al.
Published: (2025)
by: Huang, Xinzhe, et al.
Published: (2025)
Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses
by: Zheng, Xiaosen, et al.
Published: (2024)
by: Zheng, Xiaosen, et al.
Published: (2024)
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense
by: Ouyang, Yang, et al.
Published: (2025)
by: Ouyang, Yang, et al.
Published: (2025)
Directional Neural Collapse Explains Few-Shot Transfer in Self-Supervised Learning
by: Luthra, Achleshwar, et al.
Published: (2026)
by: Luthra, Achleshwar, et al.
Published: (2026)
Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks
by: Gohil, Vasudev
Published: (2025)
by: Gohil, Vasudev
Published: (2025)
Few-Shot Pattern Detection via Template Matching and Regression
by: Jo, Eunchan, et al.
Published: (2025)
by: Jo, Eunchan, et al.
Published: (2025)
Strengthening Network Intrusion Detection in IoT Environments with Self-Supervised Learning and Few Shot Learning
by: Atitallah, Safa Ben, et al.
Published: (2024)
by: Atitallah, Safa Ben, et al.
Published: (2024)
Graph-DPEP: Decomposed Plug and Ensemble Play for Few-Shot Document Relation Extraction with Graph-of-Thoughts Reasoning
by: Zhang, Tao, et al.
Published: (2024)
by: Zhang, Tao, et al.
Published: (2024)
Boosting Jailbreak Attack with Momentum
by: Zhang, Yihao, et al.
Published: (2024)
by: Zhang, Yihao, et al.
Published: (2024)
Evolving Security in LLMs: A Study of Jailbreak Attacks and Defenses
by: Shang, Zhengchun, et al.
Published: (2025)
by: Shang, Zhengchun, et al.
Published: (2025)
VisLingInstruct: Elevating Zero-Shot Learning in Multi-Modal Language Models with Autonomous Instruction Optimization
by: Zhu, Dongsheng, et al.
Published: (2024)
by: Zhu, Dongsheng, et al.
Published: (2024)
Decomposing the Basic Abilities of Large Language Models: Mitigating Cross-Task Interference in Multi-Task Instruct-Tuning
by: Wang, Bing, et al.
Published: (2026)
by: Wang, Bing, et al.
Published: (2026)
Mitigating Many-Shot Jailbreaking
by: Ackerman, Christopher M., et al.
Published: (2025)
by: Ackerman, Christopher M., et al.
Published: (2025)
Position: Universal Time Series Foundation Models Rest on a Category Error
by: Dai, Xilin, et al.
Published: (2026)
by: Dai, Xilin, et al.
Published: (2026)
Quantum Diffusion Models for Few-Shot Learning
by: Wang, Ruhan, et al.
Published: (2024)
by: Wang, Ruhan, et al.
Published: (2024)
Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing
by: Zhao, Wei, et al.
Published: (2024)
by: Zhao, Wei, et al.
Published: (2024)
Siren: A Learning-Based Multi-Turn Attack Framework for Simulating Real-World Human Jailbreak Behaviors
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
FaultDiffusion: Few-Shot Fault Time Series Generation with Diffusion Model
by: Xu, Yi, et al.
Published: (2025)
by: Xu, Yi, et al.
Published: (2025)
FusionAdapter for Few-Shot Relation Learning in Multimodal Knowledge Graphs
by: Liu, Ran, et al.
Published: (2025)
by: Liu, Ran, et al.
Published: (2025)
Subtoxic Questions: Dive Into Attitude Change of LLM's Response in Jailbreak Attempts
by: Zhang, Tianyu, et al.
Published: (2024)
by: Zhang, Tianyu, et al.
Published: (2024)
ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
by: Park, Yein, et al.
Published: (2025)
by: Park, Yein, et al.
Published: (2025)
Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs
by: Li, Jianan, et al.
Published: (2026)
by: Li, Jianan, et al.
Published: (2026)
Data Retrieval with Importance Weights for Few-Shot Imitation Learning
by: Xie, Amber, et al.
Published: (2025)
by: Xie, Amber, et al.
Published: (2025)
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
by: Wu, Yuanwei, et al.
Published: (2023)
by: Wu, Yuanwei, et al.
Published: (2023)
An experimental approach on Few Shot Class Incremental Learning
by: Adam, Marinela
Published: (2025)
by: Adam, Marinela
Published: (2025)
Uncovering the Persuasive Fingerprint of LLMs in Jailbreaking Attacks
by: Noughabi, Havva Alizadeh, et al.
Published: (2025)
by: Noughabi, Havva Alizadeh, et al.
Published: (2025)
Few-Shot Class-Incremental Learning with Prior Knowledge
by: Jiang, Wenhao, et al.
Published: (2024)
by: Jiang, Wenhao, et al.
Published: (2024)
A Strong Baseline for Molecular Few-Shot Learning
by: Formont, Philippe, et al.
Published: (2024)
by: Formont, Philippe, et al.
Published: (2024)
NTK-Guided Few-Shot Class Incremental Learning
by: Liu, Jingren, et al.
Published: (2024)
by: Liu, Jingren, et al.
Published: (2024)
FlipAttack: Jailbreak LLMs via Flipping
by: Liu, Yue, et al.
Published: (2024)
by: Liu, Yue, et al.
Published: (2024)
Efficient LLM-Jailbreaking via Multimodal-LLM Jailbreak
by: Ji, Haoxuan, et al.
Published: (2024)
by: Ji, Haoxuan, et al.
Published: (2024)
Similar Items
-
Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs
by: Yoon, Sangyeon, et al.
Published: (2026) -
SafeBehavior: Simulating Human-Like Multistage Reasoning to Mitigate Jailbreak Attacks in Large Language Models
by: Zhao, Qinjian, et al.
Published: (2025) -
Involuntary Jailbreak: On Self-Prompting Attacks
by: Guo, Yangyang, et al.
Published: (2025) -
Modality-Collaborative Low-Rank Decomposers for Few-Shot Video Domain Adaptation
by: Wanyan, Yuyang, et al.
Published: (2025) -
Enhancing Heterogeneous Knowledge Graph Completion with a Novel GAT-based Approach
by: Wei, Wanxu, et al.
Published: (2024)