CTFExplorer: Evaluating LLM Offensive Agents Through Multi-Target Web CTF Benchmarking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rani, Nanda, Milner, Kimberly, Shao, Minghao, Udeshi, Meet, Xi, Haoran, Putrevu, Venkata Sai Charan, Aggarwal, Saksham, Shukla, Sandeep K., Krishnamurthy, Prashanth, Khorrami, Farshad, Shafique, Muhammad, Karri, Ramesh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark
von: Shao, Minghao, et al.
Veröffentlicht: (2025)
von: Shao, Minghao, et al.
Veröffentlicht: (2025)
SaMOSA: Sandbox for Malware Orchestration and Side-Channel Analysis
von: Udeshi, Meet, et al.
Veröffentlicht: (2025)
von: Udeshi, Meet, et al.
Veröffentlicht: (2025)
D-CIPHER: Dynamic Collaborative Intelligent Multi-Agent System with Planner and Heterogeneous Executors for Offensive Security
von: Udeshi, Meet, et al.
Veröffentlicht: (2025)
von: Udeshi, Meet, et al.
Veröffentlicht: (2025)
AI In Cybersecurity Education -- Scalable Agentic CTF Design Principles and Educational Outcomes
von: Xi, Haoran, et al.
Veröffentlicht: (2026)
von: Xi, Haoran, et al.
Veröffentlicht: (2026)
Safeguarding LLMs Against Misuse and AI-Driven Malware Using Steganographic Canaries
von: Raz, Md, et al.
Veröffentlicht: (2026)
von: Raz, Md, et al.
Veröffentlicht: (2026)
CRAKEN: Cybersecurity LLM Agent with Knowledge-Based Execution
von: Shao, Minghao, et al.
Veröffentlicht: (2025)
von: Shao, Minghao, et al.
Veröffentlicht: (2025)
Binary Diff Summarization using Large Language Models
von: Udeshi, Meet, et al.
Veröffentlicht: (2025)
von: Udeshi, Meet, et al.
Veröffentlicht: (2025)
NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
von: Shao, Minghao, et al.
Veröffentlicht: (2024)
von: Shao, Minghao, et al.
Veröffentlicht: (2024)
SHIELD: A Host-Independent Framework for Ransomware Detection using Deep Filesystem Features
von: Raz, Md, et al.
Veröffentlicht: (2025)
von: Raz, Md, et al.
Veröffentlicht: (2025)
Ransomware 3.0: Self-Composing and LLM-Orchestrated
von: Raz, Md, et al.
Veröffentlicht: (2025)
von: Raz, Md, et al.
Veröffentlicht: (2025)
REMaQE: Reverse Engineering Math Equations from Executables
von: Udeshi, Meet, et al.
Veröffentlicht: (2023)
von: Udeshi, Meet, et al.
Veröffentlicht: (2023)
Enabling Deep Visibility into VxWorks-Based Embedded Controllers in Cyber-Physical Systems for Anomaly Detection
von: Krishnamurthy, Prashanth, et al.
Veröffentlicht: (2025)
von: Krishnamurthy, Prashanth, et al.
Veröffentlicht: (2025)
SCAMPER -- Synchrophasor Covert chAnnel for Malicious and Protective ERrands
von: Krishnamurthy, Prashanth, et al.
Veröffentlicht: (2025)
von: Krishnamurthy, Prashanth, et al.
Veröffentlicht: (2025)
Real-Time Multi-Modal Subcomponent-Level Measurements for Trustworthy System Monitoring and Malware Detection
von: Khorrami, Farshad, et al.
Veröffentlicht: (2025)
von: Khorrami, Farshad, et al.
Veröffentlicht: (2025)
Tracking Real-time Anomalies in Cyber-Physical Systems Through Dynamic Behavioral Analysis
von: Krishnamurthy, Prashanth, et al.
Veröffentlicht: (2024)
von: Krishnamurthy, Prashanth, et al.
Veröffentlicht: (2024)
HiFi-CS: Towards Open Vocabulary Visual Grounding For Robotic Grasping Using Vision-Language Models
von: Bhat, Vineet, et al.
Veröffentlicht: (2024)
von: Bhat, Vineet, et al.
Veröffentlicht: (2024)
A Comprehensive Analysis of Machine Learning Based File Trap Selection Methods to Detect Crypto Ransomware
von: Putrevu, Mohan Anand, et al.
Veröffentlicht: (2024)
von: Putrevu, Mohan Anand, et al.
Veröffentlicht: (2024)
MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping
von: Bhat, Vineet, et al.
Veröffentlicht: (2025)
von: Bhat, Vineet, et al.
Veröffentlicht: (2025)
SENTAUR: Security EnhaNced Trojan Assessment Using LLMs Against Undesirable Revisions
von: Bhandari, Jitendra, et al.
Veröffentlicht: (2024)
von: Bhandari, Jitendra, et al.
Veröffentlicht: (2024)
EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities
von: Abramovich, Talor, et al.
Veröffentlicht: (2024)
von: Abramovich, Talor, et al.
Veröffentlicht: (2024)
Grounding LLMs For Robot Task Planning Using Closed-loop State Feedback
von: Bhat, Vineet, et al.
Veröffentlicht: (2024)
von: Bhat, Vineet, et al.
Veröffentlicht: (2024)
3D CAVLA: Leveraging Depth and 3D Context to Generalize Vision Language Action Models for Unseen Tasks
von: Bhat, Vineet, et al.
Veröffentlicht: (2025)
von: Bhat, Vineet, et al.
Veröffentlicht: (2025)
Grounding Large Language Models for Robot Task Planning Using Closed‐Loop State Feedback
von: Vineet Bhat, et al.
Veröffentlicht: (2025)
von: Vineet Bhat, et al.
Veröffentlicht: (2025)
PrefixLLM: LLM-aided Prefix Circuit Design
von: Xiao, Weihua, et al.
Veröffentlicht: (2024)
von: Xiao, Weihua, et al.
Veröffentlicht: (2024)
RESCORE: LLM-Driven Simulation Recovery in Control Systems Research Papers
von: Bhat, Vineet, et al.
Veröffentlicht: (2026)
von: Bhat, Vineet, et al.
Veröffentlicht: (2026)
Data-Efficient System Identification via Lipschitz Neural Networks
von: Wei, Shiqing, et al.
Veröffentlicht: (2024)
von: Wei, Shiqing, et al.
Veröffentlicht: (2024)
RAZER: Robust Accelerated Zero-Shot 3D Open-Vocabulary Panoptic Reconstruction with Spatio-Temporal Aggregation
von: Patel, Naman, et al.
Veröffentlicht: (2025)
von: Patel, Naman, et al.
Veröffentlicht: (2025)
Confidence-Aware Safe and Stable Control of Control-Affine Systems
von: Wei, Shiqing, et al.
Veröffentlicht: (2024)
von: Wei, Shiqing, et al.
Veröffentlicht: (2024)
Combining Switching Mechanism with Re-Initialization and Anomaly Detection for Resiliency of Cyber-Physical Systems
von: Fu, Hao, et al.
Veröffentlicht: (2024)
von: Fu, Hao, et al.
Veröffentlicht: (2024)
Prescribed-Time Stability Properties of Interconnected Systems
von: Krishnamurthy, Prashanth, et al.
Veröffentlicht: (2024)
von: Krishnamurthy, Prashanth, et al.
Veröffentlicht: (2024)
Learning a Better Control Barrier Function Under Uncertain Dynamics
von: Dai, Bolun, et al.
Veröffentlicht: (2023)
von: Dai, Bolun, et al.
Veröffentlicht: (2023)
Robust Neural Lyapunov Control for Nonlinear Systems With Quadratically Bounded Disturbances
von: Shiqing Wei, et al.
Veröffentlicht: (2025)
von: Shiqing Wei, et al.
Veröffentlicht: (2025)
A Control Barrier Function-Constrained Model Predictive Control Framework for Safe Reinforcement Learning
von: Kaypak, Ali Umut, et al.
Veröffentlicht: (2026)
von: Kaypak, Ali Umut, et al.
Veröffentlicht: (2026)
OffRAMPS: An FPGA-based Intermediary for Analysis and Modification of Additive Manufacturing Control Systems
von: Blocklove, Jason, et al.
Veröffentlicht: (2024)
von: Blocklove, Jason, et al.
Veröffentlicht: (2024)
An Empirical Evaluation of LLMs for Solving Offensive Security Challenges
von: Shao, Minghao, et al.
Veröffentlicht: (2024)
von: Shao, Minghao, et al.
Veröffentlicht: (2024)
Sailing Through Point Clouds: Safe Navigation Using Point Cloud Based Control Barrier Functions
von: Dai, Bolun, et al.
Veröffentlicht: (2024)
von: Dai, Bolun, et al.
Veröffentlicht: (2024)
Out-of-Distribution Detection with Overlap Index
von: Fu, Hao, et al.
Veröffentlicht: (2024)
von: Fu, Hao, et al.
Veröffentlicht: (2024)
An Upper Bound for the Distribution Overlap Index and Its Applications
von: Fu, Hao, et al.
Veröffentlicht: (2022)
von: Fu, Hao, et al.
Veröffentlicht: (2022)
Differentiable Optimization Based Time-Varying Control Barrier Functions for Dynamic Obstacle Avoidance
von: Dai, Bolun, et al.
Veröffentlicht: (2023)
von: Dai, Bolun, et al.
Veröffentlicht: (2023)
CLIPScope: Enhancing Zero-Shot OOD Detection with Bayesian Scoring
von: Fu, Hao, et al.
Veröffentlicht: (2024)
von: Fu, Hao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark
von: Shao, Minghao, et al.
Veröffentlicht: (2025) -
SaMOSA: Sandbox for Malware Orchestration and Side-Channel Analysis
von: Udeshi, Meet, et al.
Veröffentlicht: (2025) -
D-CIPHER: Dynamic Collaborative Intelligent Multi-Agent System with Planner and Heterogeneous Executors for Offensive Security
von: Udeshi, Meet, et al.
Veröffentlicht: (2025) -
AI In Cybersecurity Education -- Scalable Agentic CTF Design Principles and Educational Outcomes
von: Xi, Haoran, et al.
Veröffentlicht: (2026) -
Safeguarding LLMs Against Misuse and AI-Driven Malware Using Steganographic Canaries
von: Raz, Md, et al.
Veröffentlicht: (2026)