Jailbreak Distillation: Renewable Safety Benchmarking
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jingyu, Elgohary, Ahmed, Wang, Xiawei, Iftekhar, A S M, Magooda, Ahmed, Van Durme, Benjamin, Khashabi, Daniel, Jackson, Kyle |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
by: Zhang, Jingyu, et al.
Published: (2024)
by: Zhang, Jingyu, et al.
Published: (2024)
SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
by: Hu, Qi, et al.
Published: (2026)
by: Hu, Qi, et al.
Published: (2026)
SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning
by: Zhou, Kaiwen, et al.
Published: (2025)
by: Zhou, Kaiwen, et al.
Published: (2025)
SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs
by: Xia, Hongfei, et al.
Published: (2025)
by: Xia, Hongfei, et al.
Published: (2025)
BackportBench: A Multilingual Benchmark for Automated Backporting of Patches
by: Zhong, Zhiqing, et al.
Published: (2025)
by: Zhong, Zhiqing, et al.
Published: (2025)
NESSiE: The Necessary Safety Benchmark -- Identifying Errors that should not Exist
by: Bertram, Johannes, et al.
Published: (2026)
by: Bertram, Johannes, et al.
Published: (2026)
Large Language Models Cannot Reliably Detect Vulnerabilities in JavaScript: The First Systematic Benchmark and Evaluation
by: Fei, Qingyuan, et al.
Published: (2025)
by: Fei, Qingyuan, et al.
Published: (2025)
Compositional Jailbreaking: An Empirical Analysis of Mutator Chain Interactions in Aligned LLMs
by: Bugnot, Reinelle Jan, et al.
Published: (2026)
by: Bugnot, Reinelle Jan, et al.
Published: (2026)
ProSec: Fortifying Code LLMs with Proactive Security Alignment
by: Xu, Xiangzhe, et al.
Published: (2024)
by: Xu, Xiangzhe, et al.
Published: (2024)
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
by: Chu, Junjie, et al.
Published: (2026)
by: Chu, Junjie, et al.
Published: (2026)
DecipherGuard: Understanding and Deciphering Jailbreak Prompts for a Safer Deployment of Intelligent Software Systems
by: Yang, Rui, et al.
Published: (2025)
by: Yang, Rui, et al.
Published: (2025)
Exposing and Defending Membership Leakage in Vulnerability Prediction Models
by: Liao, Yihan, et al.
Published: (2025)
by: Liao, Yihan, et al.
Published: (2025)
Beyond Imprecise Distance Metrics: Trace-Guided Directed Greybox Fuzzing via LLM-Predicted Call Stacks
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
HardTaint: Production-Run Dynamic Taint Analysis via Selective Hardware Tracing
by: Zhang, Yiyu, et al.
Published: (2024)
by: Zhang, Yiyu, et al.
Published: (2024)
Static Deadlock Detection for Rust Programs
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
The Secrets Must Not Flow: Scaling Security Verification to Large Codebases (extended version)
by: Arquint, Linard, et al.
Published: (2025)
by: Arquint, Linard, et al.
Published: (2025)
AutoTestForge: A Multidimensional Automated Testing Framework for Natural Language Processing Models
by: Xing, Hengrui, et al.
Published: (2025)
by: Xing, Hengrui, et al.
Published: (2025)
Argus: Reorchestrating Static Analysis via a Multi-Agent Ensemble for Full-Chain Security Vulnerability Detection
by: Liang, Zi, et al.
Published: (2026)
by: Liang, Zi, et al.
Published: (2026)
Extracting Protocol Format as State Machine via Controlled Static Loop Analysis
by: Shi, Qingkai, et al.
Published: (2023)
by: Shi, Qingkai, et al.
Published: (2023)
YASA: Scalable Multi-Language Taint Analysis on the Unified AST at Ant Group
by: Wang, Yayi, et al.
Published: (2026)
by: Wang, Yayi, et al.
Published: (2026)
SmartInv: Multimodal Learning for Smart Contract Invariant Inference
by: Wang, Sally Junsong, et al.
Published: (2024)
by: Wang, Sally Junsong, et al.
Published: (2024)
QLCoder: A Query Synthesizer For Static Analysis of Security Vulnerabilities
by: Wang, Claire, et al.
Published: (2025)
by: Wang, Claire, et al.
Published: (2025)
Translating C To Rust: Lessons from a User Study
by: Li, Ruishi, et al.
Published: (2024)
by: Li, Ruishi, et al.
Published: (2024)
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
by: Wang, Yanlin, et al.
Published: (2026)
by: Wang, Yanlin, et al.
Published: (2026)
Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
by: Chen, Xuan, et al.
Published: (2026)
by: Chen, Xuan, et al.
Published: (2026)
iResolveX: Multi-Layered Indirect Call Resolution via Static Reasoning and Learning-Augmented Refinement
by: Santra, Monika, et al.
Published: (2026)
by: Santra, Monika, et al.
Published: (2026)
IssueGuard: Real-Time Secret Leak Prevention Tool for GitHub Issue Reports
by: Rahman, Md Nafiu, et al.
Published: (2026)
by: Rahman, Md Nafiu, et al.
Published: (2026)
HALURust: Exploiting Hallucinations of Large Language Models to Detect Vulnerabilities in Rust
by: Luo, Yu, et al.
Published: (2025)
by: Luo, Yu, et al.
Published: (2025)
Usability as a Weapon: Attacking the Safety of LLM-Based Code Generation via Usability Requirements
by: Li, Yue, et al.
Published: (2026)
by: Li, Yue, et al.
Published: (2026)
Detecting Protracted Vulnerabilities in Open Source Projects
by: Sridharkumar, Arjun, et al.
Published: (2026)
by: Sridharkumar, Arjun, et al.
Published: (2026)
Managing Security Evidence in Safety-Critical Organizations
by: Mohamad, Mazen, et al.
Published: (2024)
by: Mohamad, Mazen, et al.
Published: (2024)
Online Safety Analysis for LLMs: a Benchmark, an Assessment, and a Path Forward
by: Xie, Xuan, et al.
Published: (2024)
by: Xie, Xuan, et al.
Published: (2024)
Teaching an Old LLM Secure Coding: Localized Preference Optimization on Distilled Preferences
by: Hasan, Mohammad Saqib, et al.
Published: (2025)
by: Hasan, Mohammad Saqib, et al.
Published: (2025)
CNT: Safety-oriented Function Reuse across LLMs via Cross-Model Neuron Transfer
by: Zhao, Yue, et al.
Published: (2026)
by: Zhao, Yue, et al.
Published: (2026)
Evaluating Nova 2.0 Lite model under Amazon's Frontier Model Safety Framework
by: Krishna, Satyapriya, et al.
Published: (2026)
by: Krishna, Satyapriya, et al.
Published: (2026)
Demystifying Invariant Effectiveness for Securing Smart Contracts
by: Chen, Zhiyang, et al.
Published: (2024)
by: Chen, Zhiyang, et al.
Published: (2024)
AC4: Algebraic Computation Checker for Circuit Constraints in ZKPs
by: Yang, Qizhe, et al.
Published: (2024)
by: Yang, Qizhe, et al.
Published: (2024)
IRIS: LLM-Assisted Static Analysis for Detecting Security Vulnerabilities
by: Li, Ziyang, et al.
Published: (2024)
by: Li, Ziyang, et al.
Published: (2024)
Arguzz: Testing zkVMs for Soundness and Completeness Bugs
by: Hochrainer, Christoph, et al.
Published: (2025)
by: Hochrainer, Christoph, et al.
Published: (2025)
Yaksha-Prashna: Understanding eBPF Bytecode Network Function Behavior
by: Singh, Animesh, et al.
Published: (2026)
by: Singh, Animesh, et al.
Published: (2026)
Similar Items
-
Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
by: Zhang, Jingyu, et al.
Published: (2024) -
SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
by: Hu, Qi, et al.
Published: (2026) -
SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning
by: Zhou, Kaiwen, et al.
Published: (2025) -
SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs
by: Xia, Hongfei, et al.
Published: (2025) -
BackportBench: A Multilingual Benchmark for Automated Backporting of Patches
by: Zhong, Zhiqing, et al.
Published: (2025)