Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Merves, Tyler H., Conaway, Michael H., Escobar, Joseph M., Otal, Hakan T., Tatar, Unal |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
by: Jia, Xiao
Published: (2026)
by: Jia, Xiao
Published: (2026)
Multi-Agent Synergy-Driven Iterative Visual Narrative Synthesis
by: Xi, Wang, et al.
Published: (2025)
by: Xi, Wang, et al.
Published: (2025)
Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
by: Alpay, Faruk, et al.
Published: (2025)
by: Alpay, Faruk, et al.
Published: (2025)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
by: Dang, Kieu, et al.
Published: (2025)
by: Dang, Kieu, et al.
Published: (2025)
ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization
by: Khan, Omer Jauhar
Published: (2025)
by: Khan, Omer Jauhar
Published: (2025)
Performance Evaluation of Sentiment Analysis on Text and Emoji Data Using End-to-End, Transfer Learning, Distributed and Explainable AI Models
by: Velampalli, Sirisha, et al.
Published: (2025)
by: Velampalli, Sirisha, et al.
Published: (2025)
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
by: Huo, Dongjie, et al.
Published: (2026)
by: Huo, Dongjie, et al.
Published: (2026)
LegalGuardian: A Privacy-Preserving Framework for Secure Integration of Large Language Models in Legal Practice
by: Demir, M. Mikail, et al.
Published: (2025)
by: Demir, M. Mikail, et al.
Published: (2025)
Federated Learning in Adversarial Environments: Testbed Design and Poisoning Resilience in Cybersecurity
by: Huang, Hao Jian, et al.
Published: (2024)
by: Huang, Hao Jian, et al.
Published: (2024)
When the Agent Is the Adversary: Architectural Requirements for Agentic AI Containment After the April 2026 Frontier Model Escape
by: Mitchell, Richard Joseph
Published: (2026)
by: Mitchell, Richard Joseph
Published: (2026)
SocialX: A Modular Platform for Multi-Source Big Data Research in Indonesia
by: Saputra, Muhammad Apriandito Arya, et al.
Published: (2026)
by: Saputra, Muhammad Apriandito Arya, et al.
Published: (2026)
Enhancing Mathematical Problem Solving in LLMs through Execution-Driven Reasoning Augmentation
by: Basarkar, Aditya, et al.
Published: (2026)
by: Basarkar, Aditya, et al.
Published: (2026)
MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional
by: Cruz, Roberto, et al.
Published: (2026)
by: Cruz, Roberto, et al.
Published: (2026)
LLM-Assisted Crisis Management: Building Advanced LLM Platforms for Effective Emergency Response and Public Collaboration
by: Otal, Hakan T., et al.
Published: (2024)
by: Otal, Hakan T., et al.
Published: (2024)
Adaptive Multi-Stage Patent Claim Generation with Unified Quality Assessment
by: Liang, Chen-Wei, et al.
Published: (2026)
by: Liang, Chen-Wei, et al.
Published: (2026)
HumanMCP: A Human-Like Query Dataset for Evaluating MCP Tool Retrieval Performance
by: Laddha, Shubh, et al.
Published: (2025)
by: Laddha, Shubh, et al.
Published: (2025)
Hardening x402: PII-Safe Agentic Payments via Pre-Execution Metadata Filtering
by: Stantchev, Vladimir
Published: (2026)
by: Stantchev, Vladimir
Published: (2026)
MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation
by: Qi, Dekang, et al.
Published: (2026)
by: Qi, Dekang, et al.
Published: (2026)
Multi-Agent Object Detection Framework Based on Raspberry Pi YOLO Detector and Slack-Ollama Natural Language Interface
by: Kalušev, Vladimir, et al.
Published: (2026)
by: Kalušev, Vladimir, et al.
Published: (2026)
Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
by: Li, Zonghang, et al.
Published: (2025)
by: Li, Zonghang, et al.
Published: (2025)
Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems
by: Berman, Shmuel, et al.
Published: (2024)
by: Berman, Shmuel, et al.
Published: (2024)
ARISE: Agentic Rubric-Guided Iterative Survey Engine for Automated Scholarly Paper Generation
by: Wang, Zi, et al.
Published: (2025)
by: Wang, Zi, et al.
Published: (2025)
A Multi-Agent Framework for Medical AI: Leveraging Fine-Tuned GPT, LLaMA, and DeepSeek R1 for Evidence-Based and Bias-Aware Clinical Query Processing
by: Nourmohammadi, Naeimeh, et al.
Published: (2026)
by: Nourmohammadi, Naeimeh, et al.
Published: (2026)
Evolutionary Algorithms Approach For Search Based On Semantic Document Similarity
by: Muniyappa, Chandrashekar, et al.
Published: (2025)
by: Muniyappa, Chandrashekar, et al.
Published: (2025)
Gradient Atoms: Unsupervised Discovery, Attribution and Steering of Model Behaviors via Sparse Decomposition of Training Gradients
by: Rosser, J
Published: (2026)
by: Rosser, J
Published: (2026)
Tool Receipts, Not Zero-Knowledge Proofs: Practical Hallucination Detection for AI Agents
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs
by: Hill, Brennen, et al.
Published: (2025)
by: Hill, Brennen, et al.
Published: (2025)
A V2X-based Privacy Preserving Federated Measuring and Learning System
by: Alekszejenkó, Levente, et al.
Published: (2024)
by: Alekszejenkó, Levente, et al.
Published: (2024)
LLM Honeypot: Leveraging Large Language Models as Advanced Interactive Honeypot Systems
by: Otal, Hakan T., et al.
Published: (2024)
by: Otal, Hakan T., et al.
Published: (2024)
AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks
by: Maben, Leander Melroy, et al.
Published: (2025)
by: Maben, Leander Melroy, et al.
Published: (2025)
On measuring grounding and generalizing grounding problems
by: Quigley, Daniel, et al.
Published: (2025)
by: Quigley, Daniel, et al.
Published: (2025)
AgentBreeder: Mitigating the AI Safety Risks of Multi-Agent Scaffolds via Self-Improvement
by: Rosser, J, et al.
Published: (2025)
by: Rosser, J, et al.
Published: (2025)
BreakFun: Jailbreaking LLMs via Schema Exploitation
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
Exploiting Web Search Tools of AI Agents for Data Exfiltration
by: Rall, Dennis, et al.
Published: (2025)
by: Rall, Dennis, et al.
Published: (2025)
Project Synapse: A Hierarchical Multi-Agent Framework with Hybrid Memory for Autonomous Resolution of Last-Mile Delivery Disruptions
by: Yadav, Arin Gopalan, et al.
Published: (2026)
by: Yadav, Arin Gopalan, et al.
Published: (2026)
Vertical Federated Graph Neural Network for Recommender System
by: Mai, Peihua, et al.
Published: (2023)
by: Mai, Peihua, et al.
Published: (2023)
DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
by: Wang, Zhen, et al.
Published: (2025)
by: Wang, Zhen, et al.
Published: (2025)
RACAS: Controlling Diverse Robots With a Single Agentic System
by: Ashley, Dylan R., et al.
Published: (2026)
by: Ashley, Dylan R., et al.
Published: (2026)
Geist in the Machine: Simulating Recognition and Inner Dialogue in AI-Mediated Teaching and Research
by: Magee, Liam
Published: (2026)
by: Magee, Liam
Published: (2026)
MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
by: Wang, Yihao, et al.
Published: (2026)
by: Wang, Yihao, et al.
Published: (2026)
Similar Items
-
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
by: Jia, Xiao
Published: (2026) -
Multi-Agent Synergy-Driven Iterative Visual Narrative Synthesis
by: Xi, Wang, et al.
Published: (2025) -
Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
by: Alpay, Faruk, et al.
Published: (2025) -
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
by: Dang, Kieu, et al.
Published: (2025) -
ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization
by: Khan, Omer Jauhar
Published: (2025)