Beyond Static Sandboxing: Learned Capability Governance for Autonomous AI Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Sidik, Bronislav, Rokach, Lior |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework
by: Nathanson, Samuel, et al.
Published: (2025)
by: Nathanson, Samuel, et al.
Published: (2025)
RADEP: A Resilient Adaptive Defense Framework Against Model Extraction Attacks
by: Chakraborty, Amit, et al.
Published: (2025)
by: Chakraborty, Amit, et al.
Published: (2025)
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
by: Dawson, Ads, et al.
Published: (2025)
by: Dawson, Ads, et al.
Published: (2025)
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
by: Young, Richard J., et al.
Published: (2026)
by: Young, Richard J., et al.
Published: (2026)
Governance Architecture for Autonomous Agent Systems: Threats, Framework, and Engineering Practice
by: Ge, Yuxu
Published: (2026)
by: Ge, Yuxu
Published: (2026)
Binary-30K: A Heterogeneous Dataset for Deep Learning in Binary Analysis and Malware Detection
by: Bommarito II, Michael J.
Published: (2025)
by: Bommarito II, Michael J.
Published: (2025)
Mitigating Disparate Impact of Differentially Private Learning through Bounded Adaptive Clipping
by: Zhao, Linzh, et al.
Published: (2025)
by: Zhao, Linzh, et al.
Published: (2025)
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
by: Othman, Refat
Published: (2026)
by: Othman, Refat
Published: (2026)
Operationalizing Cybersecurity Governance for Mitigation Planning with Attack-Path Modeling and Reinforcement Learning
by: Huff, Philip, et al.
Published: (2026)
by: Huff, Philip, et al.
Published: (2026)
Static Attribution of Android Residential Proxy Malware Using Graph Kernels
by: Clark, Peter, et al.
Published: (2026)
by: Clark, Peter, et al.
Published: (2026)
Dr. Jekyll and Mr. Hyde: Two Faces of LLMs
by: Collu, Matteo Gioele, et al.
Published: (2023)
by: Collu, Matteo Gioele, et al.
Published: (2023)
Evaluating Query Efficiency and Accuracy of Transfer Learning-based Model Extraction Attack in Federated Learning
by: Ahamed, Sayyed Farid, et al.
Published: (2025)
by: Ahamed, Sayyed Farid, et al.
Published: (2025)
Provable Repair of Deep Neural Network Defects by Preimage Synthesis and Property Refinement
by: Ma, Jianan, et al.
Published: (2025)
by: Ma, Jianan, et al.
Published: (2025)
MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents
by: Gowda, Ishrith
Published: (2026)
by: Gowda, Ishrith
Published: (2026)
Privately Fine-Tuned LLMs Preserve Temporal Dynamics in Tabular Data
by: Rosenblatt, Lucas, et al.
Published: (2026)
by: Rosenblatt, Lucas, et al.
Published: (2026)
SAND: A Self-supervised and Adaptive NAS-Driven Framework for Hardware Trojan Detection
by: Pan, Zhixin, et al.
Published: (2025)
by: Pan, Zhixin, et al.
Published: (2025)
The Automation Advantage in AI Red Teaming
by: Mulla, Rob, et al.
Published: (2025)
by: Mulla, Rob, et al.
Published: (2025)
KYA: A Framework-Agnostic Trust Layer for Autonomous Systems with Verifiable Provenance and Hierarchical Policy Composition
by: Quadri, Kolawole
Published: (2026)
by: Quadri, Kolawole
Published: (2026)
AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purification
by: Zhang, Tian, et al.
Published: (2026)
by: Zhang, Tian, et al.
Published: (2026)
A Self-Improving Architecture for Dynamic Safety in Large Language Models
by: Slater, Tyler
Published: (2025)
by: Slater, Tyler
Published: (2025)
Explainable Attention-Based LSTM Framework for Early Detection of AI-Assisted Ransomware via File System Behavioral Analysis
by: Nayak, Prabhudarshi, et al.
Published: (2026)
by: Nayak, Prabhudarshi, et al.
Published: (2026)
Non-Adaptive Adversarial Face Generation
by: Kim, Sunpill, et al.
Published: (2025)
by: Kim, Sunpill, et al.
Published: (2025)
When the Agent Is the Adversary: Architectural Requirements for Agentic AI Containment After the April 2026 Frontier Model Escape
by: Mitchell, Richard Joseph
Published: (2026)
by: Mitchell, Richard Joseph
Published: (2026)
AegisShield: Democratizing Cyber Threat Modeling with Generative AI
by: Grofsky, Matthew
Published: (2025)
by: Grofsky, Matthew
Published: (2025)
David vs. Goliath: Verifiable Agent-to-Agent Jailbreaking via Reinforcement Learning
by: Nellessen, Samuel, et al.
Published: (2026)
by: Nellessen, Samuel, et al.
Published: (2026)
ILION: Deterministic Pre-Execution Safety Gates for Agentic AI Systems
by: Chitan, Florin Adrian
Published: (2026)
by: Chitan, Florin Adrian
Published: (2026)
Multilingual AI-Driven Password Strength Estimation with Similarity-Based Detection
by: Palaniappan, Nikitha M., et al.
Published: (2026)
by: Palaniappan, Nikitha M., et al.
Published: (2026)
Who Governs the Machine? A Machine Identity Governance Taxonomy (MIGT) for AI Systems Operating Across Enterprise and Geopolitical Boundaries
by: Kurtz, Andrew, et al.
Published: (2026)
by: Kurtz, Andrew, et al.
Published: (2026)
Risk-Calibrated Bayesian Streaming Intrusion Detection with SRE-Aligned Decisions
by: Youssef, Michel
Published: (2025)
by: Youssef, Michel
Published: (2025)
Right to History: A Sovereignty Kernel for Verifiable AI Agent Execution
by: Zhang, Jing
Published: (2026)
by: Zhang, Jing
Published: (2026)
Session Risk Memory (SRM): Temporal Authorization for Deterministic Pre-Execution Safety Gates
by: Chitan, Florin Adrian
Published: (2026)
by: Chitan, Florin Adrian
Published: (2026)
Evaluating the Reliability of Digital Forensic Evidence Discovered by Large Language Model: A Case Study
by: Khatiwala, Jeel Piyushkumar, et al.
Published: (2026)
by: Khatiwala, Jeel Piyushkumar, et al.
Published: (2026)
Towards Agentic Investigation of Security Alerts
by: Eilertsen, Even, et al.
Published: (2026)
by: Eilertsen, Even, et al.
Published: (2026)
Poison in the Well: Feature Embedding Disruption in Backdoor Attacks
by: Feng, Zhou, et al.
Published: (2025)
by: Feng, Zhou, et al.
Published: (2025)
Attacking interpretable NLP systems
by: Abdukhamidov, Eldor, et al.
Published: (2025)
by: Abdukhamidov, Eldor, et al.
Published: (2025)
Security Considerations for Multi-agent Systems
by: Nguyen, Tam, et al.
Published: (2026)
by: Nguyen, Tam, et al.
Published: (2026)
A Practical Approach to Formal Methods: An Eclipse Integrated Development Environment (IDE) for Security Protocols
by: Garcia, Rémi, et al.
Published: (2024)
by: Garcia, Rémi, et al.
Published: (2024)
Can AI Lower the Barrier to Cybersecurity? A Human-Centered Mixed-Methods Study of Novice CTF Learning
by: Schachner, Cathrin, et al.
Published: (2026)
by: Schachner, Cathrin, et al.
Published: (2026)
Sensitivity Uncertainty Alignment in Large Language Models
by: Hiremath, Prakul Sunil, et al.
Published: (2026)
by: Hiremath, Prakul Sunil, et al.
Published: (2026)
VOLTRON: Detecting Unknown Malware Using Graph-Based Zero-Shot Learning
by: Akdeniz, M. Tahir, et al.
Published: (2025)
by: Akdeniz, M. Tahir, et al.
Published: (2025)
Similar Items
-
AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework
by: Nathanson, Samuel, et al.
Published: (2025) -
RADEP: A Resilient Adaptive Defense Framework Against Model Extraction Attacks
by: Chakraborty, Amit, et al.
Published: (2025) -
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
by: Dawson, Ads, et al.
Published: (2025) -
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
by: Young, Richard J., et al.
Published: (2026) -
Governance Architecture for Autonomous Agent Systems: Threats, Framework, and Engineering Practice
by: Ge, Yuxu
Published: (2026)