Design Principles for the Construction of a Benchmark Evaluating Security Operation Capabilities of Multi-agent AI Systems
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Cai, Yicheng, DeStefano, Mitchell John, Dong, Guodong, Handa, Pulkit, Liu, Peng, Singhal, Tejas, Tseng, Peiyu, White, Winston Jen |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
LLM Scalability Risk for Agentic-AI and Model Supply Chain Security
par: Ahi, Kiarash, et autres
Publié: (2026)
par: Ahi, Kiarash, et autres
Publié: (2026)
When the Agent Is the Adversary: Architectural Requirements for Agentic AI Containment After the April 2026 Frontier Model Escape
par: Mitchell, Richard Joseph
Publié: (2026)
par: Mitchell, Richard Joseph
Publié: (2026)
AI Agentic workflows and Enterprise APIs: Adapting API architectures for the age of AI agents
par: Tupe, Vaibhav, et autres
Publié: (2025)
par: Tupe, Vaibhav, et autres
Publié: (2025)
Security Attack and Defense Strategies for Autonomous Agent Frameworks: A Layered Review with OpenClaw as a Case Study
par: Xu, Luyao, et autres
Publié: (2026)
par: Xu, Luyao, et autres
Publié: (2026)
Security Considerations for Multi-agent Systems
par: Nguyen, Tam, et autres
Publié: (2026)
par: Nguyen, Tam, et autres
Publié: (2026)
The Importance of Out-of-Band Metadata for Safe Autonomous Agents: The Redpanda Agentic Data Plane
par: Akidau, Tyler, et autres
Publié: (2026)
par: Akidau, Tyler, et autres
Publié: (2026)
Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security
par: Qi, Jinhu, et autres
Publié: (2026)
par: Qi, Jinhu, et autres
Publié: (2026)
Bounded Autonomy for Enterprise AI: Typed Action Contracts and Consumer-Side Execution
par: Sohail, Sarmad, et autres
Publié: (2026)
par: Sohail, Sarmad, et autres
Publié: (2026)
An Organization-Scoped LLM Agent Runtime Architecture for Regulated Cybersecurity Operations
par: Fatouros, George, et autres
Publié: (2026)
par: Fatouros, George, et autres
Publié: (2026)
Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI
par: Qi, Jinhu, et autres
Publié: (2026)
par: Qi, Jinhu, et autres
Publié: (2026)
Physical oceanography during Walther Herwig III cruise WH211
par: Stein, Manfred
Publié: (2011)
par: Stein, Manfred
Publié: (2011)
Securing Agentic AI Systems -- A Multilayer Security Framework
par: Arora, Sunil, et autres
Publié: (2025)
par: Arora, Sunil, et autres
Publié: (2025)
Context Kubernetes: Declarative Orchestration of Enterprise Knowledge for Agentic AI Systems
par: Mouzouni, Charafeddine
Publié: (2026)
par: Mouzouni, Charafeddine
Publié: (2026)
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks
par: Merves, Tyler H., et autres
Publié: (2026)
par: Merves, Tyler H., et autres
Publié: (2026)
TaxAgent: How Large Language Model Designs Fiscal Policy
par: Wang, Jizhou, et autres
Publié: (2025)
par: Wang, Jizhou, et autres
Publié: (2025)
Designing Intelligent Enterprise Agents: A Capability-Aligned Multi-Agent Architecture
par: deVadoss, John
Publié: (2026)
par: deVadoss, John
Publié: (2026)
Nutrients measured in surface water during POSEIDON cruise POS211
par: OMEX Project Members, et autres
Publié: (2004)
par: OMEX Project Members, et autres
Publié: (2004)
Safe and Policy-Compliant Multi-Agent Orchestration for Enterprise AI
par: Pasupuleti, Vinil, et autres
Publié: (2026)
par: Pasupuleti, Vinil, et autres
Publié: (2026)
Agent Identity URI Scheme: Topology-Independent Naming and Capability-Based Discovery for Multi-Agent Systems
par: Rodriguez Jr, Roland R.
Publié: (2026)
par: Rodriguez Jr, Roland R.
Publié: (2026)
Physical oceanography during DISCOVERY cruise D211
par: JGOFS
Publié: (2013)
par: JGOFS
Publié: (2013)
LLM Powered Social Digital Twins: A Framework for Simulating Population Behavioral Response to Policy Interventions
par: Koaik, Fatima, et autres
Publié: (2026)
par: Koaik, Fatima, et autres
Publié: (2026)
An Agentic Multi-Agent Architecture for Cybersecurity Risk Management
par: Gupta, Ravish, et autres
Publié: (2026)
par: Gupta, Ravish, et autres
Publié: (2026)
Quantigence: A Multi-Agent AI Framework for Quantum Security Research
par: Alquwayfili, Abdulmalik
Publié: (2025)
par: Alquwayfili, Abdulmalik
Publié: (2025)
ILION: Deterministic Pre-Execution Safety Gates for Agentic AI Systems
par: Chitan, Florin Adrian
Publié: (2026)
par: Chitan, Florin Adrian
Publié: (2026)
Privacy-preserving and reward-based mechanisms of proof of engagement
par: Montanari, Matteo Marco, et autres
Publié: (2025)
par: Montanari, Matteo Marco, et autres
Publié: (2025)
Session Risk Memory (SRM): Temporal Authorization for Deterministic Pre-Execution Safety Gates
par: Chitan, Florin Adrian
Publié: (2026)
par: Chitan, Florin Adrian
Publié: (2026)
Research on Security Enhancement Methods for Adversarial Robust Large Language Model Intelligent Agents for Medical Decision-Making Tasks
par: Hu, Saisai
Publié: (2026)
par: Hu, Saisai
Publié: (2026)
PRIMA: Operational Patterns for Resilient Multi-Agent Research with Verifiable Identity and Convergent Feedback
par: Annapureddy, Sasank
Publié: (2026)
par: Annapureddy, Sasank
Publié: (2026)
Identity Management for Agentic AI: The new frontier of authorization, authentication, and security for an AI agent world
par: South, Tobin, et autres
Publié: (2025)
par: South, Tobin, et autres
Publié: (2025)
Dimethylsulphide measured in surface water during POSEIDON cruise POS211
par: OMEX Project Members, et autres
Publié: (2004)
par: OMEX Project Members, et autres
Publié: (2004)
Adaptive Episode Length Adjustment for Multi-agent Reinforcement Learning
par: Yoo, Byunghyun, et autres
Publié: (2025)
par: Yoo, Byunghyun, et autres
Publié: (2025)
Dissolved methylamines measured in surface water during POSEIDON cruise POS211
par: OMEX Project Members, et autres
Publié: (2004)
par: OMEX Project Members, et autres
Publié: (2004)
Exploring Robust Multi-Agent Workflows for Environmental Data Management
par: Guan, Boyuan, et autres
Publié: (2026)
par: Guan, Boyuan, et autres
Publié: (2026)
Requirements for Aligned, Dynamic Resolution of Conflicts in Operational Constraints
par: Jones, Steven J., et autres
Publié: (2025)
par: Jones, Steven J., et autres
Publié: (2025)
Methane and pCO2 measured in air during POSEIDON cruise POS211
par: OMEX Project Members, et autres
Publié: (2004)
par: OMEX Project Members, et autres
Publié: (2004)
Methane and pCO2 measured in surface water during POSEIDON cruise POS211
par: OMEX Project Members, et autres
Publié: (2004)
par: OMEX Project Members, et autres
Publié: (2004)
From Idea to CAD: A Language Model-Driven Multi-Agent System for Collaborative Design
par: Ocker, Felix, et autres
Publié: (2025)
par: Ocker, Felix, et autres
Publié: (2025)
HBEE: Human Behavioral Entropy Engine -- Pre-Registered Multi-Agent LLM Simulation of Peer-Suspicion-Based Detection Inversion
par: Ferrel, Vickson
Publié: (2026)
par: Ferrel, Vickson
Publié: (2026)
Depth-Dependent Indirect Prompt Injection in Tool-Calling ReAct Agents: Injection Depth, Payload Framing, and Turn-Budget Sensitivity
par: Rashidi, Mohammadreza
Publié: (2026)
par: Rashidi, Mohammadreza
Publié: (2026)
Refute-or-Promote: An Adversarial Stage-Gated Multi-Agent Review Methodology for High-Precision LLM-Assisted Defect Discovery
par: Agarwal, Abhinav
Publié: (2026)
par: Agarwal, Abhinav
Publié: (2026)
Documents similaires
-
LLM Scalability Risk for Agentic-AI and Model Supply Chain Security
par: Ahi, Kiarash, et autres
Publié: (2026) -
When the Agent Is the Adversary: Architectural Requirements for Agentic AI Containment After the April 2026 Frontier Model Escape
par: Mitchell, Richard Joseph
Publié: (2026) -
AI Agentic workflows and Enterprise APIs: Adapting API architectures for the age of AI agents
par: Tupe, Vaibhav, et autres
Publié: (2025) -
Security Attack and Defense Strategies for Autonomous Agent Frameworks: A Layered Review with OpenClaw as a Case Study
par: Xu, Luyao, et autres
Publié: (2026) -
Security Considerations for Multi-agent Systems
par: Nguyen, Tam, et autres
Publié: (2026)