NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
Fuente:
arXiv
Saved in:
| Main Authors: | Shao, Minghao, Jancheska, Sofija, Udeshi, Meet, Dolan-Gavitt, Brendan, Xi, Haoran, Milner, Kimberly, Chen, Boyuan, Yin, Max, Garg, Siddharth, Krishnamurthy, Prashanth, Khorrami, Farshad, Karri, Ramesh, Shafique, Muhammad |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Empirical Evaluation of LLMs for Solving Offensive Security Challenges
by: Shao, Minghao, et al.
Published: (2024)
by: Shao, Minghao, et al.
Published: (2024)
CTFExplorer: Evaluating LLM Offensive Agents Through Multi-Target Web CTF Benchmarking
by: Rani, Nanda, et al.
Published: (2026)
by: Rani, Nanda, et al.
Published: (2026)
AI In Cybersecurity Education -- Scalable Agentic CTF Design Principles and Educational Outcomes
by: Xi, Haoran, et al.
Published: (2026)
by: Xi, Haoran, et al.
Published: (2026)
D-CIPHER: Dynamic Collaborative Intelligent Multi-Agent System with Planner and Heterogeneous Executors for Offensive Security
by: Udeshi, Meet, et al.
Published: (2025)
by: Udeshi, Meet, et al.
Published: (2025)
Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark
by: Shao, Minghao, et al.
Published: (2025)
by: Shao, Minghao, et al.
Published: (2025)
EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities
by: Abramovich, Talor, et al.
Published: (2024)
by: Abramovich, Talor, et al.
Published: (2024)
REMaQE: Reverse Engineering Math Equations from Executables
by: Udeshi, Meet, et al.
Published: (2023)
by: Udeshi, Meet, et al.
Published: (2023)
CRAKEN: Cybersecurity LLM Agent with Knowledge-Based Execution
by: Shao, Minghao, et al.
Published: (2025)
by: Shao, Minghao, et al.
Published: (2025)
SaMOSA: Sandbox for Malware Orchestration and Side-Channel Analysis
by: Udeshi, Meet, et al.
Published: (2025)
by: Udeshi, Meet, et al.
Published: (2025)
Ransomware 3.0: Self-Composing and LLM-Orchestrated
by: Raz, Md, et al.
Published: (2025)
by: Raz, Md, et al.
Published: (2025)
Safeguarding LLMs Against Misuse and AI-Driven Malware Using Steganographic Canaries
by: Raz, Md, et al.
Published: (2026)
by: Raz, Md, et al.
Published: (2026)
Enabling Deep Visibility into VxWorks-Based Embedded Controllers in Cyber-Physical Systems for Anomaly Detection
by: Krishnamurthy, Prashanth, et al.
Published: (2025)
by: Krishnamurthy, Prashanth, et al.
Published: (2025)
SCAMPER -- Synchrophasor Covert chAnnel for Malicious and Protective ERrands
by: Krishnamurthy, Prashanth, et al.
Published: (2025)
by: Krishnamurthy, Prashanth, et al.
Published: (2025)
Real-Time Multi-Modal Subcomponent-Level Measurements for Trustworthy System Monitoring and Malware Detection
by: Khorrami, Farshad, et al.
Published: (2025)
by: Khorrami, Farshad, et al.
Published: (2025)
From Trace to Line: LLM Agent for Real-World OSS Vulnerability Localization
by: Xi, Haoran, et al.
Published: (2025)
by: Xi, Haoran, et al.
Published: (2025)
Binary Diff Summarization using Large Language Models
by: Udeshi, Meet, et al.
Published: (2025)
by: Udeshi, Meet, et al.
Published: (2025)
Tracking Real-time Anomalies in Cyber-Physical Systems Through Dynamic Behavioral Analysis
by: Krishnamurthy, Prashanth, et al.
Published: (2024)
by: Krishnamurthy, Prashanth, et al.
Published: (2024)
HiFi-CS: Towards Open Vocabulary Visual Grounding For Robotic Grasping Using Vision-Language Models
by: Bhat, Vineet, et al.
Published: (2024)
by: Bhat, Vineet, et al.
Published: (2024)
Model Cascading for Code: A Cascaded Black-Box Multi-Model Framework for Cost-Efficient Code Completion with Self-Testing
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
Out-of-Distribution Detection with Overlap Index
by: Fu, Hao, et al.
Published: (2024)
by: Fu, Hao, et al.
Published: (2024)
An Upper Bound for the Distribution Overlap Index and Its Applications
by: Fu, Hao, et al.
Published: (2022)
by: Fu, Hao, et al.
Published: (2022)
SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding
by: Ghazanfari, Sara, et al.
Published: (2026)
by: Ghazanfari, Sara, et al.
Published: (2026)
GroundCount: Grounding Vision-Language Models with Object Detection for Mitigating Counting Hallucinations
by: Chen, Boyuan, et al.
Published: (2026)
by: Chen, Boyuan, et al.
Published: (2026)
MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping
by: Bhat, Vineet, et al.
Published: (2025)
by: Bhat, Vineet, et al.
Published: (2025)
SENTAUR: Security EnhaNced Trojan Assessment Using LLMs Against Undesirable Revisions
by: Bhandari, Jitendra, et al.
Published: (2024)
by: Bhandari, Jitendra, et al.
Published: (2024)
Surgical Repair of Insecure Code Generation in LLMs
by: Sandoval, Gustavo, et al.
Published: (2026)
by: Sandoval, Gustavo, et al.
Published: (2026)
Grounding LLMs For Robot Task Planning Using Closed-loop State Feedback
by: Bhat, Vineet, et al.
Published: (2024)
by: Bhat, Vineet, et al.
Published: (2024)
3D CAVLA: Leveraging Depth and 3D Context to Generalize Vision Language Action Models for Unseen Tasks
by: Bhat, Vineet, et al.
Published: (2025)
by: Bhat, Vineet, et al.
Published: (2025)
Grounding Large Language Models for Robot Task Planning Using Closed‐Loop State Feedback
by: Vineet Bhat, et al.
Published: (2025)
by: Vineet Bhat, et al.
Published: (2025)
EMMA: Efficient Visual Alignment in Multi-Modal LLMs
by: Ghazanfari, Sara, et al.
Published: (2024)
by: Ghazanfari, Sara, et al.
Published: (2024)
LipSim: A Provably Robust Perceptual Similarity Metric
by: Ghazanfari, Sara, et al.
Published: (2023)
by: Ghazanfari, Sara, et al.
Published: (2023)
SHIELD: A Host-Independent Framework for Ransomware Detection using Deep Filesystem Features
by: Raz, Md, et al.
Published: (2025)
by: Raz, Md, et al.
Published: (2025)
RESCORE: LLM-Driven Simulation Recovery in Control Systems Research Papers
by: Bhat, Vineet, et al.
Published: (2026)
by: Bhat, Vineet, et al.
Published: (2026)
Data-Efficient System Identification via Lipschitz Neural Networks
by: Wei, Shiqing, et al.
Published: (2024)
by: Wei, Shiqing, et al.
Published: (2024)
RAZER: Robust Accelerated Zero-Shot 3D Open-Vocabulary Panoptic Reconstruction with Spatio-Temporal Aggregation
by: Patel, Naman, et al.
Published: (2025)
by: Patel, Naman, et al.
Published: (2025)
Confidence-Aware Safe and Stable Control of Control-Affine Systems
by: Wei, Shiqing, et al.
Published: (2024)
by: Wei, Shiqing, et al.
Published: (2024)
Combining Switching Mechanism with Re-Initialization and Anomaly Detection for Resiliency of Cyber-Physical Systems
by: Fu, Hao, et al.
Published: (2024)
by: Fu, Hao, et al.
Published: (2024)
Prescribed-Time Stability Properties of Interconnected Systems
by: Krishnamurthy, Prashanth, et al.
Published: (2024)
by: Krishnamurthy, Prashanth, et al.
Published: (2024)
Learning a Better Control Barrier Function Under Uncertain Dynamics
by: Dai, Bolun, et al.
Published: (2023)
by: Dai, Bolun, et al.
Published: (2023)
Robust Neural Lyapunov Control for Nonlinear Systems With Quadratically Bounded Disturbances
by: Shiqing Wei, et al.
Published: (2025)
by: Shiqing Wei, et al.
Published: (2025)
Similar Items
-
An Empirical Evaluation of LLMs for Solving Offensive Security Challenges
by: Shao, Minghao, et al.
Published: (2024) -
CTFExplorer: Evaluating LLM Offensive Agents Through Multi-Target Web CTF Benchmarking
by: Rani, Nanda, et al.
Published: (2026) -
AI In Cybersecurity Education -- Scalable Agentic CTF Design Principles and Educational Outcomes
by: Xi, Haoran, et al.
Published: (2026) -
D-CIPHER: Dynamic Collaborative Intelligent Multi-Agent System with Planner and Heterogeneous Executors for Offensive Security
by: Udeshi, Meet, et al.
Published: (2025) -
Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark
by: Shao, Minghao, et al.
Published: (2025)