When Skills Don't Help: A Negative Result on Procedural Knowledge for Tool-Grounded Agents in Offensive Cybersecurity
Fuente:
arXiv
Saved in:
| Main Authors: | Chacko, Samuel Jacob, Hugglestone, James, Islam, Chashi Mahiul, Liu, Xiuwen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
STRIATUM-CTF: A Protocol-Driven Agentic Framework for General-Purpose CTF Solving
by: Hugglestone, James, et al.
Published: (2026)
by: Hugglestone, James, et al.
Published: (2026)
Adversarial Attacks on Large Language Models Using Regularized Relaxation
by: Chacko, Samuel Jacob, et al.
Published: (2024)
by: Chacko, Samuel Jacob, et al.
Published: (2024)
DeepSeek on a Trip: Inducing Targeted Visual Hallucinations via Representation Vulnerabilities
by: Islam, Chashi Mahiul, et al.
Published: (2025)
by: Islam, Chashi Mahiul, et al.
Published: (2025)
Mechanistic Understandings of Representation Vulnerabilities and Engineering Robust Vision Transformers
by: Islam, Chashi Mahiul, et al.
Published: (2025)
by: Islam, Chashi Mahiul, et al.
Published: (2025)
Bike3S: A Tool for Bike Sharing Systems Simulation
by: Fernández, Alberto, et al.
Published: (2024)
by: Fernández, Alberto, et al.
Published: (2024)
SHM-Agents: A Generalist-Specialist Integrated Agent System for Structural Health Monitoring
by: Bao, Yuequan, et al.
Published: (2026)
by: Bao, Yuequan, et al.
Published: (2026)
PestMA: LLM-based Multi-Agent System for Informed Pest Management
by: Shi, Hongrui, et al.
Published: (2025)
by: Shi, Hongrui, et al.
Published: (2025)
Spatial-ViLT: Enhancing Visual Spatial Reasoning through Multi-Task Learning
by: Islam, Chashi Mahiul, et al.
Published: (2025)
by: Islam, Chashi Mahiul, et al.
Published: (2025)
Agreement Technologies for Coordination in Smart Cities
by: Billhardt, Holger, et al.
Published: (2024)
by: Billhardt, Holger, et al.
Published: (2024)
Autono: A ReAct-Based Highly Robust Autonomous Agent Framework
by: Wu, Zihao
Published: (2025)
by: Wu, Zihao
Published: (2025)
Centrally Coordinated Multi-Agent Reinforcement Learning for Power Grid Topology Control
by: de Mol, Barbera, et al.
Published: (2025)
by: de Mol, Barbera, et al.
Published: (2025)
From Idea to CAD: A Language Model-Driven Multi-Agent System for Collaborative Design
by: Ocker, Felix, et al.
Published: (2025)
by: Ocker, Felix, et al.
Published: (2025)
InfluenceNet: AI Models for Banzhaf and Shapley Value Prediction
by: Kempinski, Benjamin, et al.
Published: (2025)
by: Kempinski, Benjamin, et al.
Published: (2025)
Smart and Efficient IoT-Based Irrigation System Design: Utilizing a Hybrid Agent-Based and System Dynamics Approach
by: Pargo, Taha Ahmadi, et al.
Published: (2025)
by: Pargo, Taha Ahmadi, et al.
Published: (2025)
Exploring Multi-Agent Reinforcement Learning for Unrelated Parallel Machine Scheduling
by: Zampella, Maria, et al.
Published: (2024)
by: Zampella, Maria, et al.
Published: (2024)
ROTATE: Regret-driven Open-ended Training for Ad Hoc Teamwork
by: Wang, Caroline, et al.
Published: (2025)
by: Wang, Caroline, et al.
Published: (2025)
Who is Introducing the Failure? Automatically Attributing Failures of Multi-Agent Systems via Spectrum Analysis
by: Ge, Yu, et al.
Published: (2025)
by: Ge, Yu, et al.
Published: (2025)
SwarmFoam: An OpenFOAM Multi-Agent System Based on Multiple Types of Large Language Models
by: Yang, Chunwei, et al.
Published: (2026)
by: Yang, Chunwei, et al.
Published: (2026)
CS-Guide: Leveraging LLMs and Student Reflections to Provide Frequent, Scalable Academic Monitoring Feedback to Computer Science Students
by: Chacko, Samuel Jacob, et al.
Published: (2025)
by: Chacko, Samuel Jacob, et al.
Published: (2025)
Don't Retrieve, Navigate: Distilling Enterprise Knowledge into Navigable Agent Skills for QA and RAG
by: Sun, Yiqun, et al.
Published: (2026)
by: Sun, Yiqun, et al.
Published: (2026)
When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines
by: Maryanskyy, Artem
Published: (2026)
by: Maryanskyy, Artem
Published: (2026)
SkillOps: Managing LLM Agent Skill Libraries as Self-Maintaining Software Ecosystems
by: Pu, Hongji, et al.
Published: (2026)
by: Pu, Hongji, et al.
Published: (2026)
Learning Approximate Nash Equilibria in Cooperative Multi-Agent Reinforcement Learning via Mean-Field Subsampling
by: Anand, Emile, et al.
Published: (2026)
by: Anand, Emile, et al.
Published: (2026)
Prompt Engineering Guidance for Conceptual Agent-based Model Extraction using Large Language Models
by: Khatami, Siamak, et al.
Published: (2024)
by: Khatami, Siamak, et al.
Published: (2024)
Bayesian Social Deduction with Graph-Informed Language Models
by: Rahimirad, Shahab, et al.
Published: (2025)
by: Rahimirad, Shahab, et al.
Published: (2025)
GSAR: Typed Grounding for Hallucination Detection and Recovery in Multi-Agent LLMs
by: Kamelhar, Federico A.
Published: (2026)
by: Kamelhar, Federico A.
Published: (2026)
Policy Cards: Machine-Readable Runtime Governance for Autonomous AI Agents
by: Mavračić, Juraj
Published: (2025)
by: Mavračić, Juraj
Published: (2025)
Optimal Sizing and Control of a Grid-Connected Battery in a Stacked Revenue Model Including an Energy Community
by: Pocola, Tudor Octavian, et al.
Published: (2025)
by: Pocola, Tudor Octavian, et al.
Published: (2025)
Enhancing Multi-Criteria Decision Analysis with AI: Integrating Analytic Hierarchy Process and GPT-4 for Automated Decision Support
by: Svoboda, Igor, et al.
Published: (2024)
by: Svoboda, Igor, et al.
Published: (2024)
Slipstream: Trajectory-Grounded Compaction Validation for Long-Horizon Agents
by: Chen, Zhuofu, et al.
Published: (2026)
by: Chen, Zhuofu, et al.
Published: (2026)
Tool-RoCo: An Agent-as-Tool Self-organization Large Language Model Benchmark in Multi-robot Cooperation
by: Zhang, Ke, et al.
Published: (2025)
by: Zhang, Ke, et al.
Published: (2025)
BudgetMLAgent: A Cost-Effective LLM Multi-Agent system for Automating Machine Learning Tasks
by: Gandhi, Shubham, et al.
Published: (2024)
by: Gandhi, Shubham, et al.
Published: (2024)
3MDBench: Medical Multimodal Multi-agent Dialogue Benchmark
by: Sviridov, Ivan, et al.
Published: (2025)
by: Sviridov, Ivan, et al.
Published: (2025)
A survey of air combat behavior modeling using machine learning
by: Gorton, Patrick Ribu, et al.
Published: (2024)
by: Gorton, Patrick Ribu, et al.
Published: (2024)
Factorized Deep Q-Network for Cooperative Multi-Agent Reinforcement Learning in Victim Tagging
by: Cardei, Maria Ana, et al.
Published: (2025)
by: Cardei, Maria Ana, et al.
Published: (2025)
Talk is Cheap, Communication is Hard: Dynamic Grounding Failures and Repair in Multi-Agent Negotiation
by: Yao, Yiheng, et al.
Published: (2026)
by: Yao, Yiheng, et al.
Published: (2026)
VGC-Bench: Towards Mastering Diverse Team Strategies in Competitive Pokémon
by: Angliss, Cameron, et al.
Published: (2025)
by: Angliss, Cameron, et al.
Published: (2025)
Enhancing Mathematical Problem Solving in LLMs through Execution-Driven Reasoning Augmentation
by: Basarkar, Aditya, et al.
Published: (2026)
by: Basarkar, Aditya, et al.
Published: (2026)
YETI (YET to Intervene) Proactive Interventions by Multimodal AI Agents in Augmented Reality Tasks
by: Bandyopadhyay, Saptarashmi, et al.
Published: (2025)
by: Bandyopadhyay, Saptarashmi, et al.
Published: (2025)
Designing Intelligent Enterprise Agents: A Capability-Aligned Multi-Agent Architecture
by: deVadoss, John
Published: (2026)
by: deVadoss, John
Published: (2026)
Similar Items
-
STRIATUM-CTF: A Protocol-Driven Agentic Framework for General-Purpose CTF Solving
by: Hugglestone, James, et al.
Published: (2026) -
Adversarial Attacks on Large Language Models Using Regularized Relaxation
by: Chacko, Samuel Jacob, et al.
Published: (2024) -
DeepSeek on a Trip: Inducing Targeted Visual Hallucinations via Representation Vulnerabilities
by: Islam, Chashi Mahiul, et al.
Published: (2025) -
Mechanistic Understandings of Representation Vulnerabilities and Engineering Robust Vision Transformers
by: Islam, Chashi Mahiul, et al.
Published: (2025) -
Bike3S: A Tool for Bike Sharing Systems Simulation
by: Fernández, Alberto, et al.
Published: (2024)