CTFExplorer: Evaluating LLM Offensive Agents Through Multi-Target Web CTF Benchmarking
Fuente:
arXiv
Salvato in:
| Autori principali: | Rani, Nanda, Milner, Kimberly, Shao, Minghao, Udeshi, Meet, Xi, Haoran, Putrevu, Venkata Sai Charan, Aggarwal, Saksham, Shukla, Sandeep K., Krishnamurthy, Prashanth, Khorrami, Farshad, Shafique, Muhammad, Karri, Ramesh |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark
di: Shao, Minghao, et al.
Pubblicazione: (2025)
di: Shao, Minghao, et al.
Pubblicazione: (2025)
SaMOSA: Sandbox for Malware Orchestration and Side-Channel Analysis
di: Udeshi, Meet, et al.
Pubblicazione: (2025)
di: Udeshi, Meet, et al.
Pubblicazione: (2025)
D-CIPHER: Dynamic Collaborative Intelligent Multi-Agent System with Planner and Heterogeneous Executors for Offensive Security
di: Udeshi, Meet, et al.
Pubblicazione: (2025)
di: Udeshi, Meet, et al.
Pubblicazione: (2025)
AI In Cybersecurity Education -- Scalable Agentic CTF Design Principles and Educational Outcomes
di: Xi, Haoran, et al.
Pubblicazione: (2026)
di: Xi, Haoran, et al.
Pubblicazione: (2026)
Safeguarding LLMs Against Misuse and AI-Driven Malware Using Steganographic Canaries
di: Raz, Md, et al.
Pubblicazione: (2026)
di: Raz, Md, et al.
Pubblicazione: (2026)
CRAKEN: Cybersecurity LLM Agent with Knowledge-Based Execution
di: Shao, Minghao, et al.
Pubblicazione: (2025)
di: Shao, Minghao, et al.
Pubblicazione: (2025)
Binary Diff Summarization using Large Language Models
di: Udeshi, Meet, et al.
Pubblicazione: (2025)
di: Udeshi, Meet, et al.
Pubblicazione: (2025)
NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
di: Shao, Minghao, et al.
Pubblicazione: (2024)
di: Shao, Minghao, et al.
Pubblicazione: (2024)
SHIELD: A Host-Independent Framework for Ransomware Detection using Deep Filesystem Features
di: Raz, Md, et al.
Pubblicazione: (2025)
di: Raz, Md, et al.
Pubblicazione: (2025)
Ransomware 3.0: Self-Composing and LLM-Orchestrated
di: Raz, Md, et al.
Pubblicazione: (2025)
di: Raz, Md, et al.
Pubblicazione: (2025)
REMaQE: Reverse Engineering Math Equations from Executables
di: Udeshi, Meet, et al.
Pubblicazione: (2023)
di: Udeshi, Meet, et al.
Pubblicazione: (2023)
Enabling Deep Visibility into VxWorks-Based Embedded Controllers in Cyber-Physical Systems for Anomaly Detection
di: Krishnamurthy, Prashanth, et al.
Pubblicazione: (2025)
di: Krishnamurthy, Prashanth, et al.
Pubblicazione: (2025)
SCAMPER -- Synchrophasor Covert chAnnel for Malicious and Protective ERrands
di: Krishnamurthy, Prashanth, et al.
Pubblicazione: (2025)
di: Krishnamurthy, Prashanth, et al.
Pubblicazione: (2025)
Real-Time Multi-Modal Subcomponent-Level Measurements for Trustworthy System Monitoring and Malware Detection
di: Khorrami, Farshad, et al.
Pubblicazione: (2025)
di: Khorrami, Farshad, et al.
Pubblicazione: (2025)
Tracking Real-time Anomalies in Cyber-Physical Systems Through Dynamic Behavioral Analysis
di: Krishnamurthy, Prashanth, et al.
Pubblicazione: (2024)
di: Krishnamurthy, Prashanth, et al.
Pubblicazione: (2024)
HiFi-CS: Towards Open Vocabulary Visual Grounding For Robotic Grasping Using Vision-Language Models
di: Bhat, Vineet, et al.
Pubblicazione: (2024)
di: Bhat, Vineet, et al.
Pubblicazione: (2024)
A Comprehensive Analysis of Machine Learning Based File Trap Selection Methods to Detect Crypto Ransomware
di: Putrevu, Mohan Anand, et al.
Pubblicazione: (2024)
di: Putrevu, Mohan Anand, et al.
Pubblicazione: (2024)
MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping
di: Bhat, Vineet, et al.
Pubblicazione: (2025)
di: Bhat, Vineet, et al.
Pubblicazione: (2025)
SENTAUR: Security EnhaNced Trojan Assessment Using LLMs Against Undesirable Revisions
di: Bhandari, Jitendra, et al.
Pubblicazione: (2024)
di: Bhandari, Jitendra, et al.
Pubblicazione: (2024)
EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities
di: Abramovich, Talor, et al.
Pubblicazione: (2024)
di: Abramovich, Talor, et al.
Pubblicazione: (2024)
Grounding LLMs For Robot Task Planning Using Closed-loop State Feedback
di: Bhat, Vineet, et al.
Pubblicazione: (2024)
di: Bhat, Vineet, et al.
Pubblicazione: (2024)
3D CAVLA: Leveraging Depth and 3D Context to Generalize Vision Language Action Models for Unseen Tasks
di: Bhat, Vineet, et al.
Pubblicazione: (2025)
di: Bhat, Vineet, et al.
Pubblicazione: (2025)
Grounding Large Language Models for Robot Task Planning Using Closed‐Loop State Feedback
di: Vineet Bhat, et al.
Pubblicazione: (2025)
di: Vineet Bhat, et al.
Pubblicazione: (2025)
PrefixLLM: LLM-aided Prefix Circuit Design
di: Xiao, Weihua, et al.
Pubblicazione: (2024)
di: Xiao, Weihua, et al.
Pubblicazione: (2024)
RESCORE: LLM-Driven Simulation Recovery in Control Systems Research Papers
di: Bhat, Vineet, et al.
Pubblicazione: (2026)
di: Bhat, Vineet, et al.
Pubblicazione: (2026)
Data-Efficient System Identification via Lipschitz Neural Networks
di: Wei, Shiqing, et al.
Pubblicazione: (2024)
di: Wei, Shiqing, et al.
Pubblicazione: (2024)
RAZER: Robust Accelerated Zero-Shot 3D Open-Vocabulary Panoptic Reconstruction with Spatio-Temporal Aggregation
di: Patel, Naman, et al.
Pubblicazione: (2025)
di: Patel, Naman, et al.
Pubblicazione: (2025)
Confidence-Aware Safe and Stable Control of Control-Affine Systems
di: Wei, Shiqing, et al.
Pubblicazione: (2024)
di: Wei, Shiqing, et al.
Pubblicazione: (2024)
Combining Switching Mechanism with Re-Initialization and Anomaly Detection for Resiliency of Cyber-Physical Systems
di: Fu, Hao, et al.
Pubblicazione: (2024)
di: Fu, Hao, et al.
Pubblicazione: (2024)
Prescribed-Time Stability Properties of Interconnected Systems
di: Krishnamurthy, Prashanth, et al.
Pubblicazione: (2024)
di: Krishnamurthy, Prashanth, et al.
Pubblicazione: (2024)
Learning a Better Control Barrier Function Under Uncertain Dynamics
di: Dai, Bolun, et al.
Pubblicazione: (2023)
di: Dai, Bolun, et al.
Pubblicazione: (2023)
Robust Neural Lyapunov Control for Nonlinear Systems With Quadratically Bounded Disturbances
di: Shiqing Wei, et al.
Pubblicazione: (2025)
di: Shiqing Wei, et al.
Pubblicazione: (2025)
A Control Barrier Function-Constrained Model Predictive Control Framework for Safe Reinforcement Learning
di: Kaypak, Ali Umut, et al.
Pubblicazione: (2026)
di: Kaypak, Ali Umut, et al.
Pubblicazione: (2026)
OffRAMPS: An FPGA-based Intermediary for Analysis and Modification of Additive Manufacturing Control Systems
di: Blocklove, Jason, et al.
Pubblicazione: (2024)
di: Blocklove, Jason, et al.
Pubblicazione: (2024)
An Empirical Evaluation of LLMs for Solving Offensive Security Challenges
di: Shao, Minghao, et al.
Pubblicazione: (2024)
di: Shao, Minghao, et al.
Pubblicazione: (2024)
Sailing Through Point Clouds: Safe Navigation Using Point Cloud Based Control Barrier Functions
di: Dai, Bolun, et al.
Pubblicazione: (2024)
di: Dai, Bolun, et al.
Pubblicazione: (2024)
Out-of-Distribution Detection with Overlap Index
di: Fu, Hao, et al.
Pubblicazione: (2024)
di: Fu, Hao, et al.
Pubblicazione: (2024)
An Upper Bound for the Distribution Overlap Index and Its Applications
di: Fu, Hao, et al.
Pubblicazione: (2022)
di: Fu, Hao, et al.
Pubblicazione: (2022)
Differentiable Optimization Based Time-Varying Control Barrier Functions for Dynamic Obstacle Avoidance
di: Dai, Bolun, et al.
Pubblicazione: (2023)
di: Dai, Bolun, et al.
Pubblicazione: (2023)
CLIPScope: Enhancing Zero-Shot OOD Detection with Bayesian Scoring
di: Fu, Hao, et al.
Pubblicazione: (2024)
di: Fu, Hao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark
di: Shao, Minghao, et al.
Pubblicazione: (2025) -
SaMOSA: Sandbox for Malware Orchestration and Side-Channel Analysis
di: Udeshi, Meet, et al.
Pubblicazione: (2025) -
D-CIPHER: Dynamic Collaborative Intelligent Multi-Agent System with Planner and Heterogeneous Executors for Offensive Security
di: Udeshi, Meet, et al.
Pubblicazione: (2025) -
AI In Cybersecurity Education -- Scalable Agentic CTF Design Principles and Educational Outcomes
di: Xi, Haoran, et al.
Pubblicazione: (2026) -
Safeguarding LLMs Against Misuse and AI-Driven Malware Using Steganographic Canaries
di: Raz, Md, et al.
Pubblicazione: (2026)