Sockpuppetting: Jailbreaking LLMs by Combining Prefilling with Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dotsinski, Asen, Eustratiadis, Panagiotis |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
von: Li, Ran, et al.
Veröffentlicht: (2025)
von: Li, Ran, et al.
Veröffentlicht: (2025)
Improved Techniques for Optimization-Based Jailbreaking on Large Language Models
von: Jia, Xiaojun, et al.
Veröffentlicht: (2024)
von: Jia, Xiaojun, et al.
Veröffentlicht: (2024)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
Jailbreaking LLMs via Calibration
von: Lu, Yuxuan, et al.
Veröffentlicht: (2026)
von: Lu, Yuxuan, et al.
Veröffentlicht: (2026)
A StrongREJECT for Empty Jailbreaks
von: Souly, Alexandra, et al.
Veröffentlicht: (2024)
von: Souly, Alexandra, et al.
Veröffentlicht: (2024)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
von: Mehrotra, Anay, et al.
Veröffentlicht: (2023)
von: Mehrotra, Anay, et al.
Veröffentlicht: (2023)
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
von: Struppek, Lukas, et al.
Veröffentlicht: (2026)
von: Struppek, Lukas, et al.
Veröffentlicht: (2026)
Sparse Autoencoders are Capable LLM Jailbreak Mitigators
von: Assogba, Yannick, et al.
Veröffentlicht: (2026)
von: Assogba, Yannick, et al.
Veröffentlicht: (2026)
Testing the Limits of Jailbreaking Defenses with the Purple Problem
von: Kim, Taeyoun, et al.
Veröffentlicht: (2024)
von: Kim, Taeyoun, et al.
Veröffentlicht: (2024)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
von: Rando, Javier, et al.
Veröffentlicht: (2024)
von: Rando, Javier, et al.
Veröffentlicht: (2024)
VERA: Variational Inference Framework for Jailbreaking Large Language Models
von: Lochab, Anamika, et al.
Veröffentlicht: (2025)
von: Lochab, Anamika, et al.
Veröffentlicht: (2025)
Deciphering the Chaos: Enhancing Jailbreak Attacks via Adversarial Prompt Translation
von: Li, Qizhang, et al.
Veröffentlicht: (2024)
von: Li, Qizhang, et al.
Veröffentlicht: (2024)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks
von: Hu, Hanjiang, et al.
Veröffentlicht: (2025)
von: Hu, Hanjiang, et al.
Veröffentlicht: (2025)
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
von: Liang, Buyun, et al.
Veröffentlicht: (2025)
von: Liang, Buyun, et al.
Veröffentlicht: (2025)
Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-Tuning
von: Hasan, Adib, et al.
Veröffentlicht: (2024)
von: Hasan, Adib, et al.
Veröffentlicht: (2024)
Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement
von: Kim, Heegyu, et al.
Veröffentlicht: (2024)
von: Kim, Heegyu, et al.
Veröffentlicht: (2024)
LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models
von: Lin, Shi, et al.
Veröffentlicht: (2024)
von: Lin, Shi, et al.
Veröffentlicht: (2024)
From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
RED QUEEN: Safeguarding Large Language Models against Concealed Multi-Turn Jailbreaking
von: Jiang, Yifan, et al.
Veröffentlicht: (2024)
von: Jiang, Yifan, et al.
Veröffentlicht: (2024)
Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models
von: Hu, Kai, et al.
Veröffentlicht: (2025)
von: Hu, Kai, et al.
Veröffentlicht: (2025)
Jailbreaking in the Haystack
von: Shah, Rishi Rajesh, et al.
Veröffentlicht: (2025)
von: Shah, Rishi Rajesh, et al.
Veröffentlicht: (2025)
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling
von: Ma, Avery, et al.
Veröffentlicht: (2025)
von: Ma, Avery, et al.
Veröffentlicht: (2025)
Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities
von: Geng, Jiahui, et al.
Veröffentlicht: (2025)
von: Geng, Jiahui, et al.
Veröffentlicht: (2025)
PrisonBreak: Jailbreaking Large Language Models with at Most Twenty-Five Targeted Bit-flips
von: Coalson, Zachary, et al.
Veröffentlicht: (2024)
von: Coalson, Zachary, et al.
Veröffentlicht: (2024)
TwinGate: Stateful Defense against Decompositional Jailbreaks in Untraceable Traffic via Asymmetric Contrastive Learning
von: Sun, Bowen, et al.
Veröffentlicht: (2026)
von: Sun, Bowen, et al.
Veröffentlicht: (2026)
Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?
von: Chen, Shuo, et al.
Veröffentlicht: (2024)
von: Chen, Shuo, et al.
Veröffentlicht: (2024)
Jailbreaking with Universal Multi-Prompts
von: Hsu, Yu-Ling, et al.
Veröffentlicht: (2025)
von: Hsu, Yu-Ling, et al.
Veröffentlicht: (2025)
Faster-GCG: Efficient Discrete Optimization Jailbreak Attacks against Aligned Large Language Models
von: Li, Xiao, et al.
Veröffentlicht: (2024)
von: Li, Xiao, et al.
Veröffentlicht: (2024)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
von: Li, Nathaniel, et al.
Veröffentlicht: (2024)
von: Li, Nathaniel, et al.
Veröffentlicht: (2024)
AdvPrefix: An Objective for Nuanced LLM Jailbreaks
von: Zhu, Sicheng, et al.
Veröffentlicht: (2024)
von: Zhu, Sicheng, et al.
Veröffentlicht: (2024)
Rethinking How to Evaluate Language Model Jailbreak
von: Cai, Hongyu, et al.
Veröffentlicht: (2024)
von: Cai, Hongyu, et al.
Veröffentlicht: (2024)
Low-Resource Languages Jailbreak GPT-4
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2023)
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2023)
Jailbreaking Large Language Models with Symbolic Mathematics
von: Bethany, Emet, et al.
Veröffentlicht: (2024)
von: Bethany, Emet, et al.
Veröffentlicht: (2024)
How Vulnerable Are Edge LLMs?
von: Ding, Ao, et al.
Veröffentlicht: (2026)
von: Ding, Ao, et al.
Veröffentlicht: (2026)
Universal Jailbreak Backdoors from Poisoned Human Feedback
von: Rando, Javier, et al.
Veröffentlicht: (2023)
von: Rando, Javier, et al.
Veröffentlicht: (2023)
EnJa: Ensemble Jailbreak on Large Language Models
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
UCD: Unlearning in LLMs via Contrastive Decoding
von: Suriyakumar, Vinith M., et al.
Veröffentlicht: (2025)
von: Suriyakumar, Vinith M., et al.
Veröffentlicht: (2025)
Coercing LLMs to do and reveal (almost) anything
von: Geiping, Jonas, et al.
Veröffentlicht: (2024)
von: Geiping, Jonas, et al.
Veröffentlicht: (2024)
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
von: Xiong, Alexander, et al.
Veröffentlicht: (2025)
von: Xiong, Alexander, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
von: Li, Ran, et al.
Veröffentlicht: (2025) -
Improved Techniques for Optimization-Based Jailbreaking on Large Language Models
von: Jia, Xiaojun, et al.
Veröffentlicht: (2024) -
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
von: Chu, Junjie, et al.
Veröffentlicht: (2024) -
Jailbreaking LLMs via Calibration
von: Lu, Yuxuan, et al.
Veröffentlicht: (2026) -
A StrongREJECT for Empty Jailbreaks
von: Souly, Alexandra, et al.
Veröffentlicht: (2024)