An Empirical Evaluation of LLMs for Solving Offensive Security Challenges
Fuente:
arXiv
Saved in:
| Main Authors: | Shao, Minghao, Chen, Boyuan, Jancheska, Sofija, Dolan-Gavitt, Brendan, Garg, Siddharth, Karri, Ramesh, Shafique, Muhammad |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
by: Shao, Minghao, et al.
Published: (2024)
by: Shao, Minghao, et al.
Published: (2024)
From Trace to Line: LLM Agent for Real-World OSS Vulnerability Localization
by: Xi, Haoran, et al.
Published: (2025)
by: Xi, Haoran, et al.
Published: (2025)
Surgical Repair of Insecure Code Generation in LLMs
by: Sandoval, Gustavo, et al.
Published: (2026)
by: Sandoval, Gustavo, et al.
Published: (2026)
D-CIPHER: Dynamic Collaborative Intelligent Multi-Agent System with Planner and Heterogeneous Executors for Offensive Security
by: Udeshi, Meet, et al.
Published: (2025)
by: Udeshi, Meet, et al.
Published: (2025)
MetaCipher: A Time-Persistent and Universal Multi-Agent Framework for Cipher-Based Jailbreak Attacks for LLMs
by: Chen, Boyuan, et al.
Published: (2025)
by: Chen, Boyuan, et al.
Published: (2025)
(Security) Assertions by Large Language Models
by: Kande, Rahul, et al.
Published: (2023)
by: Kande, Rahul, et al.
Published: (2023)
TrojanLoC: LLM-based Framework for RTL Trojan Localization
by: Xiao, Weihua, et al.
Published: (2025)
by: Xiao, Weihua, et al.
Published: (2025)
HarmChip: Evaluating Hardware Security Centric LLM Safety via Jailbreak Benchmarking
by: Wang, Zeng, et al.
Published: (2026)
by: Wang, Zeng, et al.
Published: (2026)
NetDeTox: Adversarial and Efficient Evasion of Hardware-Security GNNs via RL-LLM Orchestration
by: Wang, Zeng, et al.
Published: (2025)
by: Wang, Zeng, et al.
Published: (2025)
ELFuzz: Efficient Input Generation via LLM-driven Synthesis Over Fuzzer Space
by: Chen, Chuyang, et al.
Published: (2025)
by: Chen, Chuyang, et al.
Published: (2025)
Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark
by: Shao, Minghao, et al.
Published: (2025)
by: Shao, Minghao, et al.
Published: (2025)
CTFExplorer: Evaluating LLM Offensive Agents Through Multi-Target Web CTF Benchmarking
by: Rani, Nanda, et al.
Published: (2026)
by: Rani, Nanda, et al.
Published: (2026)
LLMs for Secure Hardware Design and Related Problems: Opportunities and Challenges
by: Knechtel, Johann, et al.
Published: (2026)
by: Knechtel, Johann, et al.
Published: (2026)
CRAKEN: Cybersecurity LLM Agent with Knowledge-Based Execution
by: Shao, Minghao, et al.
Published: (2025)
by: Shao, Minghao, et al.
Published: (2025)
VeriContaminated: Assessing LLM-Driven Verilog Coding for Data Contamination
by: Wang, Zeng, et al.
Published: (2025)
by: Wang, Zeng, et al.
Published: (2025)
What is the AGI in Offensive Security ?
by: Cho, Youngwoong
Published: (2026)
by: Cho, Youngwoong
Published: (2026)
TrojanGYM: A Detector-in-the-Loop LLM for Adaptive RTL Hardware Trojan Insertion
by: Sreekumar, Saideep, et al.
Published: (2026)
by: Sreekumar, Saideep, et al.
Published: (2026)
SALAD: Systematic Assessment of Machine Unlearning on LLM-Aided Hardware Design
by: Wang, Zeng, et al.
Published: (2025)
by: Wang, Zeng, et al.
Published: (2025)
LASHED: LLMs And Static Hardware Analysis for Early Detection of RTL Bugs
by: Ahmad, Baleegh, et al.
Published: (2025)
by: Ahmad, Baleegh, et al.
Published: (2025)
Model Cascading for Code: A Cascaded Black-Box Multi-Model Framework for Cost-Efficient Code Completion with Self-Testing
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
Fixing Hardware Security Bugs with Large Language Models
by: Ahmad, Baleegh, et al.
Published: (2023)
by: Ahmad, Baleegh, et al.
Published: (2023)
GroundCount: Grounding Vision-Language Models with Object Detection for Mitigating Counting Hallucinations
by: Chen, Boyuan, et al.
Published: (2026)
by: Chen, Boyuan, et al.
Published: (2026)
VeriLeaky: Navigating IP Protection vs Utility in Fine-Tuning for LLM-Driven Verilog Coding
by: Wang, Zeng, et al.
Published: (2025)
by: Wang, Zeng, et al.
Published: (2025)
ASCENT: Amplifying Power Side-Channel Resilience via Learning & Monte-Carlo Tree Search
by: Bhandari, Jitendra, et al.
Published: (2024)
by: Bhandari, Jitendra, et al.
Published: (2024)
Need for zkSpeed: Accelerating HyperPlonk for Zero-Knowledge Proofs
by: Daftardar, Alhad, et al.
Published: (2025)
by: Daftardar, Alhad, et al.
Published: (2025)
Enabling Deep Visibility into VxWorks-Based Embedded Controllers in Cyber-Physical Systems for Anomaly Detection
by: Krishnamurthy, Prashanth, et al.
Published: (2025)
by: Krishnamurthy, Prashanth, et al.
Published: (2025)
SCAMPER -- Synchrophasor Covert chAnnel for Malicious and Protective ERrands
by: Krishnamurthy, Prashanth, et al.
Published: (2025)
by: Krishnamurthy, Prashanth, et al.
Published: (2025)
Real-Time Multi-Modal Subcomponent-Level Measurements for Trustworthy System Monitoring and Malware Detection
by: Khorrami, Farshad, et al.
Published: (2025)
by: Khorrami, Farshad, et al.
Published: (2025)
SENTAUR: Security EnhaNced Trojan Assessment Using LLMs Against Undesirable Revisions
by: Bhandari, Jitendra, et al.
Published: (2024)
by: Bhandari, Jitendra, et al.
Published: (2024)
RTL-Breaker: Assessing the Security of LLMs against Backdoor Attacks on HDL Code Generation
by: Mankali, Lakshmi Likhitha, et al.
Published: (2024)
by: Mankali, Lakshmi Likhitha, et al.
Published: (2024)
Safeguarding LLMs Against Misuse and AI-Driven Malware Using Steganographic Canaries
by: Raz, Md, et al.
Published: (2026)
by: Raz, Md, et al.
Published: (2026)
LLM4PQC - Accurate and Efficient Synthesis of PQC Cores by Feedback-Driven LLMs
by: Perera, Buddhi, et al.
Published: (2026)
by: Perera, Buddhi, et al.
Published: (2026)
Security Challenges of Complex Space Applications: An Empirical Study
by: Paulik, Tomas
Published: (2024)
by: Paulik, Tomas
Published: (2024)
Offensive Security for AI Systems: Concepts, Practices, and Applications
by: Harguess, Josh, et al.
Published: (2025)
by: Harguess, Josh, et al.
Published: (2025)
Artificial Intelligence as the New Hacker: Developing Agents for Offensive Security
by: Valencia, Leroy Jacob
Published: (2024)
by: Valencia, Leroy Jacob
Published: (2024)
Offensive Robot Cybersecurity
by: Mayoral-Vilches, Víctor
Published: (2025)
by: Mayoral-Vilches, Víctor
Published: (2025)
Exploring the Robustness and Transferability of Patch-Based Adversarial Attacks in Quantized Neural Networks
by: Guesmi, Amira, et al.
Published: (2024)
by: Guesmi, Amira, et al.
Published: (2024)
LLMs in the SOC: An Empirical Study of Human-AI Collaboration in Security Operations Centres
by: Singh, Ronal, et al.
Published: (2025)
by: Singh, Ronal, et al.
Published: (2025)
Anomaly Unveiled: Securing Image Classification against Adversarial Patch Attacks
by: Chattopadhyay, Nandish, et al.
Published: (2024)
by: Chattopadhyay, Nandish, et al.
Published: (2024)
LockForge: Automating Paper-to-Code for Logic Locking with Multi-Agent Reasoning LLMs
by: Saha, Akashdeep, et al.
Published: (2025)
by: Saha, Akashdeep, et al.
Published: (2025)
Similar Items
-
NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
by: Shao, Minghao, et al.
Published: (2024) -
From Trace to Line: LLM Agent for Real-World OSS Vulnerability Localization
by: Xi, Haoran, et al.
Published: (2025) -
Surgical Repair of Insecure Code Generation in LLMs
by: Sandoval, Gustavo, et al.
Published: (2026) -
D-CIPHER: Dynamic Collaborative Intelligent Multi-Agent System with Planner and Heterogeneous Executors for Offensive Security
by: Udeshi, Meet, et al.
Published: (2025) -
MetaCipher: A Time-Persistent and Universal Multi-Agent Framework for Cipher-Based Jailbreak Attacks for LLMs
by: Chen, Boyuan, et al.
Published: (2025)