SATA: A Paradigm for LLM Jailbreak via Simple Assistive Task Linkage
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dong, Xiaoning, Hu, Wenbo, Xu, Wei, He, Tianxing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ShieldLearner: A New Paradigm for Jailbreak Attack Defense in LLMs
von: Ni, Ziyi, et al.
Veröffentlicht: (2025)
von: Ni, Ziyi, et al.
Veröffentlicht: (2025)
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
von: Xing, Wenpeng, et al.
Veröffentlicht: (2025)
von: Xing, Wenpeng, et al.
Veröffentlicht: (2025)
Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial Finetuning
von: Sel, Bilgehan, et al.
Veröffentlicht: (2026)
von: Sel, Bilgehan, et al.
Veröffentlicht: (2026)
LLM Jailbreak Detection for (Almost) Free!
von: Chen, Guorui, et al.
Veröffentlicht: (2025)
von: Chen, Guorui, et al.
Veröffentlicht: (2025)
SSG: Logit-Balanced Vocabulary Partitioning for LLM Watermarking
von: Gu, Chenxi, et al.
Veröffentlicht: (2026)
von: Gu, Chenxi, et al.
Veröffentlicht: (2026)
Geneshift: Impact of different scenario shift on Jailbreaking LLM
von: Wu, Tianyi, et al.
Veröffentlicht: (2025)
von: Wu, Tianyi, et al.
Veröffentlicht: (2025)
Beyond Jailbreaking: Auditing Contextual Privacy in LLM Agents
von: Das, Saswat, et al.
Veröffentlicht: (2025)
von: Das, Saswat, et al.
Veröffentlicht: (2025)
LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models
von: Yu, Miao, et al.
Veröffentlicht: (2024)
von: Yu, Miao, et al.
Veröffentlicht: (2024)
GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation
von: Ramesh, Govind, et al.
Veröffentlicht: (2024)
von: Ramesh, Govind, et al.
Veröffentlicht: (2024)
JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models
von: Ran, Delong, et al.
Veröffentlicht: (2024)
von: Ran, Delong, et al.
Veröffentlicht: (2024)
LLM for SoC Security: A Paradigm Shift
von: Saha, Dipayan, et al.
Veröffentlicht: (2023)
von: Saha, Dipayan, et al.
Veröffentlicht: (2023)
Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency
von: Zhao, Shiji, et al.
Veröffentlicht: (2025)
von: Zhao, Shiji, et al.
Veröffentlicht: (2025)
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
von: Li, Xirui, et al.
Veröffentlicht: (2024)
von: Li, Xirui, et al.
Veröffentlicht: (2024)
Subtoxic Questions: Dive Into Attitude Change of LLM's Response in Jailbreak Attempts
von: Zhang, Tianyu, et al.
Veröffentlicht: (2024)
von: Zhang, Tianyu, et al.
Veröffentlicht: (2024)
Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space
von: Huang, Yao, et al.
Veröffentlicht: (2025)
von: Huang, Yao, et al.
Veröffentlicht: (2025)
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models
von: Xu, Zihao, et al.
Veröffentlicht: (2024)
von: Xu, Zihao, et al.
Veröffentlicht: (2024)
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2025)
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2025)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
von: Xu, Zhao, et al.
Veröffentlicht: (2024)
von: Xu, Zhao, et al.
Veröffentlicht: (2024)
TrailBlazer: History-Guided Reinforcement Learning for Black-Box LLM Jailbreaking
von: Yoon, Sung-Hoon, et al.
Veröffentlicht: (2026)
von: Yoon, Sung-Hoon, et al.
Veröffentlicht: (2026)
How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation
von: Long, Zhuohang, et al.
Veröffentlicht: (2025)
von: Long, Zhuohang, et al.
Veröffentlicht: (2025)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
von: Liu, Fan, et al.
Veröffentlicht: (2024)
von: Liu, Fan, et al.
Veröffentlicht: (2024)
Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia
von: Shen, Guangyu, et al.
Veröffentlicht: (2024)
von: Shen, Guangyu, et al.
Veröffentlicht: (2024)
Reverse-Engineering Model Editing on Language Models
von: Sun, Zhiyu, et al.
Veröffentlicht: (2026)
von: Sun, Zhiyu, et al.
Veröffentlicht: (2026)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation
von: Schwartz, Daniel, et al.
Veröffentlicht: (2025)
von: Schwartz, Daniel, et al.
Veröffentlicht: (2025)
Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks
von: Kabir, Md Rysul, et al.
Veröffentlicht: (2026)
von: Kabir, Md Rysul, et al.
Veröffentlicht: (2026)
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
von: Ji, Wence, et al.
Veröffentlicht: (2025)
von: Ji, Wence, et al.
Veröffentlicht: (2025)
Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models
von: Yu, Yongcan, et al.
Veröffentlicht: (2025)
von: Yu, Yongcan, et al.
Veröffentlicht: (2025)
Unleashing the Unseen: Harnessing Benign Datasets for Jailbreaking Large Language Models
von: Zhao, Wei, et al.
Veröffentlicht: (2024)
von: Zhao, Wei, et al.
Veröffentlicht: (2024)
from Benign import Toxic: Jailbreaking the Language Model via Adversarial Metaphors
von: Yan, Yu, et al.
Veröffentlicht: (2025)
von: Yan, Yu, et al.
Veröffentlicht: (2025)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2024)
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2024)
SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters
von: Yang, Yan, et al.
Veröffentlicht: (2024)
von: Yang, Yan, et al.
Veröffentlicht: (2024)
Beyond the Tip of Efficiency: Uncovering the Submerged Threats of Jailbreak Attacks in Small Language Models
von: Yi, Sibo, et al.
Veröffentlicht: (2025)
von: Yi, Sibo, et al.
Veröffentlicht: (2025)
FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
von: Gong, Yichen, et al.
Veröffentlicht: (2023)
von: Gong, Yichen, et al.
Veröffentlicht: (2023)
MRJ-Agent: An Effective Jailbreak Agent for Multi-Round Dialogue
von: Wang, Fengxiang, et al.
Veröffentlicht: (2024)
von: Wang, Fengxiang, et al.
Veröffentlicht: (2024)
AdvPrefix: An Objective for Nuanced LLM Jailbreaks
von: Zhu, Sicheng, et al.
Veröffentlicht: (2024)
von: Zhu, Sicheng, et al.
Veröffentlicht: (2024)
AdaPPA: Adaptive Position Pre-Fill Jailbreak Attack Approach Targeting LLMs
von: Lv, Lijia, et al.
Veröffentlicht: (2024)
von: Lv, Lijia, et al.
Veröffentlicht: (2024)
ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models
von: Cheng, Siyang, et al.
Veröffentlicht: (2025)
von: Cheng, Siyang, et al.
Veröffentlicht: (2025)
RobustKV: Defending Large Language Models against Jailbreak Attacks via KV Eviction
von: Jiang, Tanqiu, et al.
Veröffentlicht: (2024)
von: Jiang, Tanqiu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ShieldLearner: A New Paradigm for Jailbreak Attack Defense in LLMs
von: Ni, Ziyi, et al.
Veröffentlicht: (2025) -
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
von: Xing, Wenpeng, et al.
Veröffentlicht: (2025) -
Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial Finetuning
von: Sel, Bilgehan, et al.
Veröffentlicht: (2026) -
LLM Jailbreak Detection for (Almost) Free!
von: Chen, Guorui, et al.
Veröffentlicht: (2025) -
SSG: Logit-Balanced Vocabulary Partitioning for LLM Watermarking
von: Gu, Chenxi, et al.
Veröffentlicht: (2026)