JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Chao, Patrick, Debenedetti, Edoardo, Robey, Alexander, Andriushchenko, Maksym, Croce, Francesco, Sehwag, Vikash, Dobriban, Edgar, Flammarion, Nicolas, Pappas, George J., Tramer, Florian, Hassani, Hamed, Wong, Eric |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
Jailbreaking Black Box Large Language Models in Twenty Queries
di: Chao, Patrick, et al.
Pubblicazione: (2023)
di: Chao, Patrick, et al.
Pubblicazione: (2023)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
di: Rando, Javier, et al.
Pubblicazione: (2024)
di: Rando, Javier, et al.
Pubblicazione: (2024)
SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
di: Robey, Alexander, et al.
Pubblicazione: (2023)
di: Robey, Alexander, et al.
Pubblicazione: (2023)
Jailbreaking LLM-Controlled Robots
di: Robey, Alexander, et al.
Pubblicazione: (2024)
di: Robey, Alexander, et al.
Pubblicazione: (2024)
Long Is More for Alignment: A Simple but Tough-to-Beat Baseline for Instruction Fine-Tuning
di: Zhao, Hao, et al.
Pubblicazione: (2024)
di: Zhao, Hao, et al.
Pubblicazione: (2024)
Is In-Context Learning Sufficient for Instruction Following in LLMs?
di: Zhao, Hao, et al.
Pubblicazione: (2024)
di: Zhao, Hao, et al.
Pubblicazione: (2024)
Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing
di: Ji, Jiabao, et al.
Pubblicazione: (2024)
di: Ji, Jiabao, et al.
Pubblicazione: (2024)
Exploring Memorization and Copyright Violation in Frontier LLMs: A Study of the New York Times v. OpenAI 2023 Lawsuit
di: Freeman, Joshua, et al.
Pubblicazione: (2024)
di: Freeman, Joshua, et al.
Pubblicazione: (2024)
Provable tradeoffs in adversarially robust classification
di: Dobriban, Edgar, et al.
Pubblicazione: (2020)
di: Dobriban, Edgar, et al.
Pubblicazione: (2020)
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
di: Carlini, Nicholas, et al.
Pubblicazione: (2025)
di: Carlini, Nicholas, et al.
Pubblicazione: (2025)
Adversarial Reasoning at Jailbreaking Time
di: Sabbaghi, Mahdi, et al.
Pubblicazione: (2025)
di: Sabbaghi, Mahdi, et al.
Pubblicazione: (2025)
Does Refusal Training in LLMs Generalize to the Past Tense?
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
Evading Black-box Classifiers Without Breaking Eggs
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2023)
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2023)
Adversarial Search Engine Optimization for Large Language Models
di: Nestaas, Fredrik, et al.
Pubblicazione: (2024)
di: Nestaas, Fredrik, et al.
Pubblicazione: (2024)
Conformal Inference under High-Dimensional Covariate Shifts via Likelihood-Ratio Regularization
di: Joshi, Sunay, et al.
Pubblicazione: (2025)
di: Joshi, Sunay, et al.
Pubblicazione: (2025)
OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
di: Kuntz, Thomas, et al.
Pubblicazione: (2025)
di: Kuntz, Thomas, et al.
Pubblicazione: (2025)
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
di: Nikolić, Kristina, et al.
Pubblicazione: (2025)
di: Nikolić, Kristina, et al.
Pubblicazione: (2025)
Preventing Robotic Jailbreaking via Multimodal Domain Adaptation
di: Marchiori, Francesco, et al.
Pubblicazione: (2025)
di: Marchiori, Francesco, et al.
Pubblicazione: (2025)
Universal Jailbreak Backdoors from Poisoned Human Feedback
di: Rando, Javier, et al.
Pubblicazione: (2023)
di: Rando, Javier, et al.
Pubblicazione: (2023)
HalluHard: A Hard Multi-Turn Hallucination Benchmark
di: Fan, Dongyang, et al.
Pubblicazione: (2026)
di: Fan, Dongyang, et al.
Pubblicazione: (2026)
Benchmarking Misuse Mitigation Against Covert Adversaries
di: Brown, Davis, et al.
Pubblicazione: (2025)
di: Brown, Davis, et al.
Pubblicazione: (2025)
Contextual Safety Reasoning and Grounding for Open-World Robots
di: Ravichandran, Zachary, et al.
Pubblicazione: (2026)
di: Ravichandran, Zachary, et al.
Pubblicazione: (2026)
Evaluating the Performance of Large Language Models via Debates
di: Moniri, Behrad, et al.
Pubblicazione: (2024)
di: Moniri, Behrad, et al.
Pubblicazione: (2024)
Watermarking Language Models with Error Correcting Codes
di: Chao, Patrick, et al.
Pubblicazione: (2024)
di: Chao, Patrick, et al.
Pubblicazione: (2024)
Why Do We Need Weight Decay in Modern Deep Learning?
di: D'Angelo, Francesco, et al.
Pubblicazione: (2023)
di: D'Angelo, Francesco, et al.
Pubblicazione: (2023)
Safety Guardrails for LLM-Enabled Robots
di: Ravichandran, Zachary, et al.
Pubblicazione: (2025)
di: Ravichandran, Zachary, et al.
Pubblicazione: (2025)
Adversarial Training Should Be Cast as a Non-Zero-Sum Game
di: Robey, Alexander, et al.
Pubblicazione: (2023)
di: Robey, Alexander, et al.
Pubblicazione: (2023)
Gradient-based Jailbreak Images for Multimodal Fusion Models
di: Rando, Javier, et al.
Pubblicazione: (2024)
di: Rando, Javier, et al.
Pubblicazione: (2024)
QuantSightBench: Evaluating LLM Quantitative Forecasting with Prediction Intervals
di: Qin, Jeremy, et al.
Pubblicazione: (2026)
di: Qin, Jeremy, et al.
Pubblicazione: (2026)
Jailbreaking in the Haystack
di: Shah, Rishi Rajesh, et al.
Pubblicazione: (2025)
di: Shah, Rishi Rajesh, et al.
Pubblicazione: (2025)
A Theory of Non-Linear Feature Learning with One Gradient Step in Two-Layer Neural Networks
di: Moniri, Behrad, et al.
Pubblicazione: (2023)
di: Moniri, Behrad, et al.
Pubblicazione: (2023)
Risk-Controlled Post-Processing of Decision Policies
di: Joshi, Sunay, et al.
Pubblicazione: (2026)
di: Joshi, Sunay, et al.
Pubblicazione: (2026)
MultiRisk: Multiple Risk Control via Iterative Score Thresholding
di: Joshi, Sunay, et al.
Pubblicazione: (2025)
di: Joshi, Sunay, et al.
Pubblicazione: (2025)
Chordal Sparsity for Lipschitz Constant Estimation of Deep Neural Networks
di: Xue, Anton, et al.
Pubblicazione: (2022)
di: Xue, Anton, et al.
Pubblicazione: (2022)
Measuring Non-Adversarial Reproduction of Training Data in Large Language Models
di: Aerni, Michael, et al.
Pubblicazione: (2024)
di: Aerni, Michael, et al.
Pubblicazione: (2024)
Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks
di: Hu, Hanjiang, et al.
Pubblicazione: (2025)
di: Hu, Hanjiang, et al.
Pubblicazione: (2025)
Length Optimization in Conformal Prediction
di: Kiyani, Shayan, et al.
Pubblicazione: (2024)
di: Kiyani, Shayan, et al.
Pubblicazione: (2024)
Conformal Prediction with Learned Features
di: Kiyani, Shayan, et al.
Pubblicazione: (2024)
di: Kiyani, Shayan, et al.
Pubblicazione: (2024)
Selective Induction Heads: How Transformers Select Causal Structures In Context
di: D'Angelo, Francesco, et al.
Pubblicazione: (2025)
di: D'Angelo, Francesco, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024) -
Jailbreaking Black Box Large Language Models in Twenty Queries
di: Chao, Patrick, et al.
Pubblicazione: (2023) -
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
di: Rando, Javier, et al.
Pubblicazione: (2024) -
SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
di: Robey, Alexander, et al.
Pubblicazione: (2023) -
Jailbreaking LLM-Controlled Robots
di: Robey, Alexander, et al.
Pubblicazione: (2024)