Saved in:
| Main Authors: | Li, Changyi, Lu, Pengfei, Pan, Xudong, Barez, Fazl, Yang, Min |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.07427 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly
by: Fan, Yihe, et al.
Published: (2026)
by: Fan, Yihe, et al.
Published: (2026)
PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
by: Fu, Tingchen, et al.
Published: (2024)
by: Fu, Tingchen, et al.
Published: (2024)
Large language model-powered AI systems achieve self-replication with no human intervention
by: Pan, Xudong, et al.
Published: (2025)
by: Pan, Xudong, et al.
Published: (2025)
Security Risks in Tool-Enabled AI Agents: A Systematic Analysis of Privileged Execution Environments
by: Goel, Hardik
Published: (2026)
by: Goel, Hardik
Published: (2026)
Neural Dehydration: Effective Erasure of Black-box Watermarks from DNNs with Limited Data
by: Lu, Yifan, et al.
Published: (2023)
by: Lu, Yifan, et al.
Published: (2023)
Invisible Threats from Model Context Protocol: Generating Stealthy Injection Payload via Tree-based Adaptive Search
by: Shen, Yulin, et al.
Published: (2026)
by: Shen, Yulin, et al.
Published: (2026)
LinuxArena: A Control Setting for AI Agents in Live Production Software Environments
by: Tracy, Tyler, et al.
Published: (2026)
by: Tracy, Tyler, et al.
Published: (2026)
ForesightSafety Bench: A Frontier Risk Evaluation and Governance Framework towards Safe AI
by: Tong, Haibo, et al.
Published: (2026)
by: Tong, Haibo, et al.
Published: (2026)
SentinelNet: Safeguarding Multi-Agent Collaboration Through Credit-Based Dynamic Threat Detection
by: Feng, Yang, et al.
Published: (2025)
by: Feng, Yang, et al.
Published: (2025)
StruPhantom: Evolutionary Injection Attacks on Black-Box Tabular Agents Powered by Large Language Models
by: Feng, Yang, et al.
Published: (2025)
by: Feng, Yang, et al.
Published: (2025)
BashArena: A Control Setting for Highly Privileged AI Agents
by: Kaufman, Adam, et al.
Published: (2025)
by: Kaufman, Adam, et al.
Published: (2025)
WebTrap Park: An Automated Platform for Systematic Security Evaluation of Web Agents
by: Wu, Xinyi, et al.
Published: (2026)
by: Wu, Xinyi, et al.
Published: (2026)
BELT: Old-School Backdoor Attacks can Evade the State-of-the-Art Defense with Backdoor Exclusivity Lifting
by: Qiu, Huming, et al.
Published: (2023)
by: Qiu, Huming, et al.
Published: (2023)
Quantifying CBRN Risk in Frontier Models
by: Kumar, Divyanshu, et al.
Published: (2025)
by: Kumar, Divyanshu, et al.
Published: (2025)
An AI Agent Execution Environment to Safeguard User Data
by: Stanley, Robert, et al.
Published: (2026)
by: Stanley, Robert, et al.
Published: (2026)
FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments
by: Yang, Zhi, et al.
Published: (2026)
by: Yang, Zhi, et al.
Published: (2026)
Graph in the Vault: Protecting Edge GNN Inference with Trusted Execution Environment
by: Ding, Ruyi, et al.
Published: (2025)
by: Ding, Ruyi, et al.
Published: (2025)
Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach
by: Wang, Tony T., et al.
Published: (2024)
by: Wang, Tony T., et al.
Published: (2024)
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments
by: Schnabl, Christoph, et al.
Published: (2025)
by: Schnabl, Christoph, et al.
Published: (2025)
Accelerating AI Development with Cyber Arenas
by: Cashman, William, et al.
Published: (2025)
by: Cashman, William, et al.
Published: (2025)
AutoPenBench: Benchmarking Generative Agents for Penetration Testing
by: Gioacchini, Luca, et al.
Published: (2024)
by: Gioacchini, Luca, et al.
Published: (2024)
DMind-3: A Sovereign Edge--Local--Cloud AI System with Controlled Deliberation and Correction-Based Tuning for Safe, Low-Latency Transaction Execution
by: Huang, Enhao, et al.
Published: (2026)
by: Huang, Enhao, et al.
Published: (2026)
AutoStub: Genetic Programming-Based Stub Creation for Symbolic Execution
by: Mächtle, Felix, et al.
Published: (2025)
by: Mächtle, Felix, et al.
Published: (2025)
Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw
by: Yao, Hongwei, et al.
Published: (2026)
by: Yao, Hongwei, et al.
Published: (2026)
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
by: Panfilov, Alexander, et al.
Published: (2025)
by: Panfilov, Alexander, et al.
Published: (2025)
Red-Teaming Claude Opus and ChatGPT-based Security Advisors for Trusted Execution Environments
by: Mukherjee, Kunal, et al.
Published: (2026)
by: Mukherjee, Kunal, et al.
Published: (2026)
Measuring the Permission Gate: A Stress-Test Evaluation of Claude Code's Auto Mode
by: Ji, Zimo, et al.
Published: (2026)
by: Ji, Zimo, et al.
Published: (2026)
PRISON: Unmasking the Criminal Potential of Large Language Models
by: Wu, Xinyi, et al.
Published: (2025)
by: Wu, Xinyi, et al.
Published: (2025)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
by: Kutasov, Jonathan, et al.
Published: (2025)
by: Kutasov, Jonathan, et al.
Published: (2025)
Frontier AI's Impact on the Cybersecurity Landscape
by: Potter, Yujin, et al.
Published: (2025)
by: Potter, Yujin, et al.
Published: (2025)
AutoDAN-Reasoning: Enhancing Strategies Exploration based Jailbreak Attacks with Test-Time Scaling
by: Liu, Xiaogeng, et al.
Published: (2025)
by: Liu, Xiaogeng, et al.
Published: (2025)
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
by: Li, Yige, et al.
Published: (2025)
by: Li, Yige, et al.
Published: (2025)
AutoPT: How Far Are We from the End2End Automated Web Penetration Testing?
by: Wu, Benlong, et al.
Published: (2024)
by: Wu, Benlong, et al.
Published: (2024)
Unveiling Privacy Risks in LLM Agent Memory
by: Wang, Bo, et al.
Published: (2025)
by: Wang, Bo, et al.
Published: (2025)
Statistical Proof of Execution (SPEX)
by: Dallachiesa, Michele, et al.
Published: (2025)
by: Dallachiesa, Michele, et al.
Published: (2025)
On the Feasibility of Using MultiModal LLMs to Execute AR Social Engineering Attacks
by: Bi, Ting, et al.
Published: (2025)
by: Bi, Ting, et al.
Published: (2025)
AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration
by: Zhou, Andy, et al.
Published: (2025)
by: Zhou, Andy, et al.
Published: (2025)
AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use
by: Yang, Chenglin
Published: (2026)
by: Yang, Chenglin
Published: (2026)
Multi-Agent Penetration Testing AI for the Web
by: David, Isaac, et al.
Published: (2025)
by: David, Isaac, et al.
Published: (2025)
SHIELD: An Auto-Healing Agentic Defense Framework for LLM Resource Exhaustion Attacks
by: Sivaroopan, Nirhoshan, et al.
Published: (2026)
by: Sivaroopan, Nirhoshan, et al.
Published: (2026)
Similar Items
-
CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly
by: Fan, Yihe, et al.
Published: (2026) -
PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
by: Fu, Tingchen, et al.
Published: (2024) -
Large language model-powered AI systems achieve self-replication with no human intervention
by: Pan, Xudong, et al.
Published: (2025) -
Security Risks in Tool-Enabled AI Agents: A Systematic Analysis of Privileged Execution Environments
by: Goel, Hardik
Published: (2026) -
Neural Dehydration: Effective Erasure of Black-box Watermarks from DNNs with Limited Data
by: Lu, Yifan, et al.
Published: (2023)