Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Rando, Javier, Croce, Francesco, Mitka, Kryštof, Shabalin, Stepan, Andriushchenko, Maksym, Flammarion, Nicolas, Tramèr, Florian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
Universal Jailbreak Backdoors from Poisoned Human Feedback
por: Rando, Javier, et al.
Publicado: (2023)
por: Rando, Javier, et al.
Publicado: (2023)
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
por: Chao, Patrick, et al.
Publicado: (2024)
por: Chao, Patrick, et al.
Publicado: (2024)
Gradient-based Jailbreak Images for Multimodal Fusion Models
por: Rando, Javier, et al.
Publicado: (2024)
por: Rando, Javier, et al.
Publicado: (2024)
Adversarial Perturbations Cannot Reliably Protect Artists From Generative AI
por: Hönig, Robert, et al.
Publicado: (2024)
por: Hönig, Robert, et al.
Publicado: (2024)
Is In-Context Learning Sufficient for Instruction Following in LLMs?
por: Zhao, Hao, et al.
Publicado: (2024)
por: Zhao, Hao, et al.
Publicado: (2024)
Adversarial ML Problems Are Getting Harder to Solve and to Evaluate
por: Rando, Javier, et al.
Publicado: (2025)
por: Rando, Javier, et al.
Publicado: (2025)
Privacy Backdoors: Stealing Data with Corrupted Pretrained Models
por: Feng, Shanglun, et al.
Publicado: (2024)
por: Feng, Shanglun, et al.
Publicado: (2024)
An Adversarial Perspective on Machine Unlearning for AI Safety
por: Łucki, Jakub, et al.
Publicado: (2024)
por: Łucki, Jakub, et al.
Publicado: (2024)
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
por: Ji, Wence, et al.
Publicado: (2025)
por: Ji, Wence, et al.
Publicado: (2025)
Long Is More for Alignment: A Simple but Tough-to-Beat Baseline for Instruction Fine-Tuning
por: Zhao, Hao, et al.
Publicado: (2024)
por: Zhao, Hao, et al.
Publicado: (2024)
Does Refusal Training in LLMs Generalize to the Past Tense?
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
por: Wang, Jiongxiao, et al.
Publicado: (2024)
por: Wang, Jiongxiao, et al.
Publicado: (2024)
Blind Baselines Beat Membership Inference Attacks for Foundation Models
por: Das, Debeshee, et al.
Publicado: (2024)
por: Das, Debeshee, et al.
Publicado: (2024)
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
por: Nikolić, Kristina, et al.
Publicado: (2025)
por: Nikolić, Kristina, et al.
Publicado: (2025)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
por: Panfilov, Alexander, et al.
Publicado: (2025)
por: Panfilov, Alexander, et al.
Publicado: (2025)
Let the Bees Find the Weak Spots: A Path Planning Perspective on Multi-Turn Jailbreak Attacks against LLMs
por: Liu, Yize, et al.
Publicado: (2025)
por: Liu, Yize, et al.
Publicado: (2025)
Mitigating Jailbreaks with Intent-Aware LLMs
por: Yeo, Wei Jie, et al.
Publicado: (2025)
por: Yeo, Wei Jie, et al.
Publicado: (2025)
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
por: Chen, Zhuowei, et al.
Publicado: (2025)
por: Chen, Zhuowei, et al.
Publicado: (2025)
Persistent Pre-Training Poisoning of LLMs
por: Zhang, Yiming, et al.
Publicado: (2024)
por: Zhang, Yiming, et al.
Publicado: (2024)
AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
por: Zhao, Weixiang, et al.
Publicado: (2025)
por: Zhao, Weixiang, et al.
Publicado: (2025)
MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs
por: Wen, Rui, et al.
Publicado: (2026)
por: Wen, Rui, et al.
Publicado: (2026)
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
por: Carlini, Nicholas, et al.
Publicado: (2025)
por: Carlini, Nicholas, et al.
Publicado: (2025)
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
por: Guo, Weiyang, et al.
Publicado: (2026)
por: Guo, Weiyang, et al.
Publicado: (2026)
A Simple and Efficient Jailbreak Method Exploiting LLMs' Helpfulness
por: Luo, Xuan, et al.
Publicado: (2025)
por: Luo, Xuan, et al.
Publicado: (2025)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
por: Zhang, Chiyu, et al.
Publicado: (2025)
por: Zhang, Chiyu, et al.
Publicado: (2025)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
por: Chen, Yunhao, et al.
Publicado: (2025)
por: Chen, Yunhao, et al.
Publicado: (2025)
BitBypass: A New Direction in Jailbreaking Aligned Large Language Models with Bitstream Camouflage
por: Nakka, Kalyan, et al.
Publicado: (2025)
por: Nakka, Kalyan, et al.
Publicado: (2025)
ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs
por: Zhao, Gejian, et al.
Publicado: (2025)
por: Zhao, Gejian, et al.
Publicado: (2025)
Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-Tuning
por: Hasan, Adib, et al.
Publicado: (2024)
por: Hasan, Adib, et al.
Publicado: (2024)
Evaluating the Robustness of the "Ensemble Everything Everywhere" Defense
por: Zhang, Jie, et al.
Publicado: (2024)
por: Zhang, Jie, et al.
Publicado: (2024)
One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs
por: Li, Linbao, et al.
Publicado: (2025)
por: Li, Linbao, et al.
Publicado: (2025)
Jailbreaking LLMs via Semantically Relevant Nested Scenarios with Targeted Toxic Knowledge
por: Xu, Ning, et al.
Publicado: (2025)
por: Xu, Ning, et al.
Publicado: (2025)
BadActs: A Universal Backdoor Defense in the Activation Space
por: Yi, Biao, et al.
Publicado: (2024)
por: Yi, Biao, et al.
Publicado: (2024)
TuBA: Cross-Lingual Transferability of Backdoor Attacks in LLMs with Instruction Tuning
por: He, Xuanli, et al.
Publicado: (2024)
por: He, Xuanli, et al.
Publicado: (2024)
GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis
por: Xie, Yueqi, et al.
Publicado: (2024)
por: Xie, Yueqi, et al.
Publicado: (2024)
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
por: Li, Xiang, et al.
Publicado: (2025)
por: Li, Xiang, et al.
Publicado: (2025)
PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization
por: Wang, Yidan, et al.
Publicado: (2025)
por: Wang, Yidan, et al.
Publicado: (2025)
Overlooked Safety Vulnerability in LLMs: Malicious Intelligent Optimization Algorithm Request and its Jailbreak
por: Gu, Haoran, et al.
Publicado: (2026)
por: Gu, Haoran, et al.
Publicado: (2026)
ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models
por: Cheng, Siyang, et al.
Publicado: (2025)
por: Cheng, Siyang, et al.
Publicado: (2025)
Ejemplares similares
-
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
por: Andriushchenko, Maksym, et al.
Publicado: (2024) -
Universal Jailbreak Backdoors from Poisoned Human Feedback
por: Rando, Javier, et al.
Publicado: (2023) -
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
por: Chao, Patrick, et al.
Publicado: (2024) -
Gradient-based Jailbreak Images for Multimodal Fusion Models
por: Rando, Javier, et al.
Publicado: (2024) -
Adversarial Perturbations Cannot Reliably Protect Artists From Generative AI
por: Hönig, Robert, et al.
Publicado: (2024)