Saved in:
| Main Authors: | Marchand, Rahul, Cathain, Art O, Wynne, Jerome, Giavridis, Philippos Maximos, Deverett, Sam, Wilkinson, John, Gwartz, Jason, Coppock, Harry |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.02277 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mythos and the Unverified Cage: Z3-Based Pre-Deployment Verification for Frontier-Model Sandbox Infrastructure
by: Blain, Dominik
Published: (2026)
by: Blain, Dominik
Published: (2026)
On the Differential Privacy and Interactivity of Privacy Sandbox Reports
by: Ghazi, Badih, et al.
Published: (2024)
by: Ghazi, Badih, et al.
Published: (2024)
Quantifying CBRN Risk in Frontier Models
by: Kumar, Divyanshu, et al.
Published: (2025)
by: Kumar, Divyanshu, et al.
Published: (2025)
On the Feasibility of CubeSats Application Sandboxing for Space Missions
by: Marra, Gabriele, et al.
Published: (2024)
by: Marra, Gabriele, et al.
Published: (2024)
SaMOSA: Sandbox for Malware Orchestration and Side-Channel Analysis
by: Udeshi, Meet, et al.
Published: (2025)
by: Udeshi, Meet, et al.
Published: (2025)
Sandboxing Adoption in Open Source Ecosystems
by: Alhindi, Maysara, et al.
Published: (2024)
by: Alhindi, Maysara, et al.
Published: (2024)
Scaling Up: Revisiting Mining Android Sandboxes at Scale for Malware Classification
by: Costa, Francisco, et al.
Published: (2025)
by: Costa, Francisco, et al.
Published: (2025)
Technical Report: The Need for a (Research) Sandstorm through the Privacy Sandbox
by: Beugin, Yohan, et al.
Published: (2025)
by: Beugin, Yohan, et al.
Published: (2025)
PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction
by: Garza, Leon, et al.
Published: (2025)
by: Garza, Leon, et al.
Published: (2025)
ceLLMate: Sandboxing Browser AI Agents
by: Meng, Luoxi, et al.
Published: (2025)
by: Meng, Luoxi, et al.
Published: (2025)
Jailbroken Frontier Models Retain Their Capabilities
by: Zhu, Daniel, et al.
Published: (2026)
by: Zhu, Daniel, et al.
Published: (2026)
Threadbox: Sandboxing for Modular Security
by: Alhindi, Maysara, et al.
Published: (2025)
by: Alhindi, Maysara, et al.
Published: (2025)
VirtualCrime: Evaluating Criminal Potential of Large Language Models via Sandbox Simulation
by: Tang, Yilin, et al.
Published: (2026)
by: Tang, Yilin, et al.
Published: (2026)
SandCell: Sandboxing Rust Beyond Unsafe Code
by: Zhang, Jialun, et al.
Published: (2025)
by: Zhang, Jialun, et al.
Published: (2025)
A Large Language Model Approach to Generating Bypass Rules for Malware Evasion in Analysis Sandbox
by: Sui, Zhiyong, et al.
Published: (2026)
by: Sui, Zhiyong, et al.
Published: (2026)
Okapi: Efficiently Safeguarding Speculative Data Accesses in Sandboxed Environments
by: Schmitz, Philipp, et al.
Published: (2023)
by: Schmitz, Philipp, et al.
Published: (2023)
SandboxEval: Towards Securing Test Environment for Untrusted Code
by: Rabin, Rafiqul, et al.
Published: (2025)
by: Rabin, Rafiqul, et al.
Published: (2025)
SoK: An Essential Guide For Using Malware Sandboxes In Security Applications: Challenges, Pitfalls, and Lessons Learned
by: Alrawi, Omar, et al.
Published: (2024)
by: Alrawi, Omar, et al.
Published: (2024)
Dynamic Frequency-Based Fingerprinting Attacks against Modern Sandbox Environments
by: Dipta, Debopriya Roy, et al.
Published: (2024)
by: Dipta, Debopriya Roy, et al.
Published: (2024)
Lifefin: Escaping Mempool Explosions in DAG-based BFT
by: Zhang, Jianting, et al.
Published: (2025)
by: Zhang, Jianting, et al.
Published: (2025)
LaserEscape: Detecting and Mitigating Optical Probing Attacks
by: Monfared, Saleh Khalaj, et al.
Published: (2024)
by: Monfared, Saleh Khalaj, et al.
Published: (2024)
pokiSEC: A Multi-Architecture, Containerized Ephemeral Malware Detonation Sandbox
by: Avina, Alejandro, et al.
Published: (2025)
by: Avina, Alejandro, et al.
Published: (2025)
WebGPU-SPY: Finding Fingerprints in the Sandbox through GPU Cache Attacks
by: Ferguson, Ethan, et al.
Published: (2024)
by: Ferguson, Ethan, et al.
Published: (2024)
Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
by: Anurin, Andrey, et al.
Published: (2024)
by: Anurin, Andrey, et al.
Published: (2024)
Outside the Comfort Zone: Analysing LLM Capabilities in Software Vulnerability Detection
by: Guo, Yuejun, et al.
Published: (2024)
by: Guo, Yuejun, et al.
Published: (2024)
Early Signs of Steganographic Capabilities in Frontier LLMs
by: Zolkowski, Artur, et al.
Published: (2025)
by: Zolkowski, Artur, et al.
Published: (2025)
Beyond Static Sandboxing: Learned Capability Governance for Autonomous AI Agents
by: Sidik, Bronislav, et al.
Published: (2026)
by: Sidik, Bronislav, et al.
Published: (2026)
MCP-SandboxScan: WASM-based Secure Execution and Runtime Analysis for MCP Tools
by: Tan, Zhuoran, et al.
Published: (2026)
by: Tan, Zhuoran, et al.
Published: (2026)
Beyond Jailbreak: Unveiling Risks in LLM Applications Arising from Blurred Capability Boundaries
by: Zhang, Yunyi, et al.
Published: (2025)
by: Zhang, Yunyi, et al.
Published: (2025)
Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework
by: Krishna, Satyapriya, et al.
Published: (2025)
by: Krishna, Satyapriya, et al.
Published: (2025)
Quantifying Loss Aversion in Cyber Adversaries via LLM Analysis
by: Hans, Soham, et al.
Published: (2025)
by: Hans, Soham, et al.
Published: (2025)
Evaluating LLM Generated Detection Rules in Cybersecurity
by: Bertiger, Anna, et al.
Published: (2025)
by: Bertiger, Anna, et al.
Published: (2025)
Measuring Modern Phishing Tactics: A Quantitative Study of Body Obfuscation Prevalence, Co-occurrence, and Filter Impact
by: Dalmiere, Antony, et al.
Published: (2025)
by: Dalmiere, Antony, et al.
Published: (2025)
Quantifying Association Capabilities of Large Language Models and Its Implications on Privacy Leakage
by: Shao, Hanyin, et al.
Published: (2023)
by: Shao, Hanyin, et al.
Published: (2023)
Evaluating Nova 2.0 Lite model under Amazon's Frontier Model Safety Framework
by: Krishna, Satyapriya, et al.
Published: (2026)
by: Krishna, Satyapriya, et al.
Published: (2026)
Digital Agriculture Sandbox for Collaborative Research
by: Zafar, Osama, et al.
Published: (2025)
by: Zafar, Osama, et al.
Published: (2025)
ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
by: Lee, Seunghyun, et al.
Published: (2026)
by: Lee, Seunghyun, et al.
Published: (2026)
Enhancing Linux Privilege Escalation Attack Capabilities of Local LLM Agents
by: Probst, Benjamin, et al.
Published: (2026)
by: Probst, Benjamin, et al.
Published: (2026)
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
by: Che, Zora, et al.
Published: (2025)
by: Che, Zora, et al.
Published: (2025)
Quantifying Azure RBAC Wildcard Overreach
by: Parisel, Christophe
Published: (2025)
by: Parisel, Christophe
Published: (2025)
Similar Items
-
Mythos and the Unverified Cage: Z3-Based Pre-Deployment Verification for Frontier-Model Sandbox Infrastructure
by: Blain, Dominik
Published: (2026) -
On the Differential Privacy and Interactivity of Privacy Sandbox Reports
by: Ghazi, Badih, et al.
Published: (2024) -
Quantifying CBRN Risk in Frontier Models
by: Kumar, Divyanshu, et al.
Published: (2025) -
On the Feasibility of CubeSats Application Sandboxing for Space Missions
by: Marra, Gabriele, et al.
Published: (2024) -
SaMOSA: Sandbox for Malware Orchestration and Side-Channel Analysis
by: Udeshi, Meet, et al.
Published: (2025)