StepShield: When, Not Whether to Intervene on Rogue Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Felicia, Gloria, Eniolade, Michael, He, Jinfeng, Sasindran, Zitha, Kumar, Hemant, Angati, Milan Hussain, Bandarupalli, Sandeep |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When "Correct" Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents?
by: Peng, Yibo, et al.
Published: (2025)
by: Peng, Yibo, et al.
Published: (2025)
When Specifications Meet Reality: Uncovering API Inconsistencies in Ethereum Infrastructure
by: Ma, Jie, et al.
Published: (2026)
by: Ma, Jie, et al.
Published: (2026)
When MCP Servers Attack: Taxonomy, Feasibility, and Mitigation
by: Zhao, Weibo, et al.
Published: (2025)
by: Zhao, Weibo, et al.
Published: (2025)
Take a Step Further: Understanding Page Spray in Linux Kernel Exploitation
by: Guo, Ziyi, et al.
Published: (2024)
by: Guo, Ziyi, et al.
Published: (2024)
When Security Meets Usability: An Empirical Investigation of Post-Quantum Cryptography APIs
by: Toruan, Marthin, et al.
Published: (2026)
by: Toruan, Marthin, et al.
Published: (2026)
When AI Takes the Wheel: Security Analysis of Framework-Constrained Program Generation
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems
by: Wang, Su, et al.
Published: (2026)
by: Wang, Su, et al.
Published: (2026)
FuzzAgent: Multi-Agent System for Evolutionary Library Fuzzing
by: Lyu, Yunlong, et al.
Published: (2026)
by: Lyu, Yunlong, et al.
Published: (2026)
When Labels Are Scarce: A Systematic Mapping of Label-Efficient Code Vulnerability Detection
by: Khalal, Noor, et al.
Published: (2026)
by: Khalal, Noor, et al.
Published: (2026)
Beyond Function-Level Analysis: Context-Aware Reasoning for Inter-Procedural Vulnerability Detection
by: Li, Yikun, et al.
Published: (2026)
by: Li, Yikun, et al.
Published: (2026)
SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration
by: Guo, Zihan, et al.
Published: (2026)
by: Guo, Zihan, et al.
Published: (2026)
Microservice Vulnerability Analysis: A Literature Review with Empirical Insights
by: Jayalath, Raveen Kanishka, et al.
Published: (2024)
by: Jayalath, Raveen Kanishka, et al.
Published: (2024)
Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications
by: Lu, Xiaoyue, et al.
Published: (2026)
by: Lu, Xiaoyue, et al.
Published: (2026)
From LLMs to Agents: A Comparative Evaluation of LLMs and LLM-based Agents in Security Patch Detection
by: Han, Junxiao, et al.
Published: (2025)
by: Han, Junxiao, et al.
Published: (2025)
AgentGuard: A Multi-Agent Framework for Robust Package Confusion Detection via Hybrid Search and Metadata-Content Fusion
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
Multi-Agent Taint Specification Extraction for Vulnerability Detection
by: Ghebremichael, Jonah, et al.
Published: (2026)
by: Ghebremichael, Jonah, et al.
Published: (2026)
Semantics-Aligned, Curriculum-Driven, and Reasoning-Enhanced Vulnerability Repair Framework
by: Yang, Chengran, et al.
Published: (2025)
by: Yang, Chengran, et al.
Published: (2025)
CHASE: LLM Agents for Dissecting Malicious PyPI Packages
by: Toda, Takaaki, et al.
Published: (2026)
by: Toda, Takaaki, et al.
Published: (2026)
Identifying Adversary Tactics and Techniques in Malware Binaries with an LLM Agent
by: Xuan, Zhou, et al.
Published: (2026)
by: Xuan, Zhou, et al.
Published: (2026)
LLM Agents for Automated Web Vulnerability Reproduction: Are We There Yet?
by: Liu, Bin, et al.
Published: (2025)
by: Liu, Bin, et al.
Published: (2025)
LLMs as Firmware Experts: A Runtime-Grown Tree-of-Agents Framework
by: Zhang, Xiangrui, et al.
Published: (2025)
by: Zhang, Xiangrui, et al.
Published: (2025)
ARGUS: Defending LLM Agents Against Context-Aware Prompt Injection
by: Weng, Shihao, et al.
Published: (2026)
by: Weng, Shihao, et al.
Published: (2026)
Exploiting LLM Agent Supply Chains via Payload-less Skills
by: Liu, Xinyu, et al.
Published: (2026)
by: Liu, Xinyu, et al.
Published: (2026)
MARD: A Multi-Agent Framework for Robust Android Malware Detection
by: Zeng, Xueying, et al.
Published: (2026)
by: Zeng, Xueying, et al.
Published: (2026)
Coverage-Guided Multi-Agent Harness Generation for Java Library Fuzzing
by: Loose, Nils, et al.
Published: (2026)
by: Loose, Nils, et al.
Published: (2026)
SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
by: Hu, Qi, et al.
Published: (2026)
by: Hu, Qi, et al.
Published: (2026)
What Makes a Good LLM Agent for Real-world Penetration Testing?
by: Deng, Gelei, et al.
Published: (2026)
by: Deng, Gelei, et al.
Published: (2026)
Multi-Agent Collaborative Fuzzing with Continuous Reflection for Smart Contracts Vulnerability Detection
by: Chen, Jie, et al.
Published: (2025)
by: Chen, Jie, et al.
Published: (2025)
Think Broad, Act Narrow: CWE Identification with Multi-Agent Large Language Models
by: Sayagh, Mohammed, et al.
Published: (2025)
by: Sayagh, Mohammed, et al.
Published: (2025)
Clawdrain: Exploiting Tool-Calling Chains for Stealthy Token Exhaustion in OpenClaw Agents
by: Dong, Ben, et al.
Published: (2026)
by: Dong, Ben, et al.
Published: (2026)
HarnessAgent: Scaling Automatic Fuzzing Harness Construction with Tool-Augmented LLM Pipelines
by: Yang, Kang, et al.
Published: (2025)
by: Yang, Kang, et al.
Published: (2025)
ChainFuzzer: Greybox Fuzzing for Workflow-Level Multi-Tool Vulnerabilities in LLM Agents
by: Wu, Jiangrong, et al.
Published: (2026)
by: Wu, Jiangrong, et al.
Published: (2026)
Security Is Relative: Training-Free Vulnerability Detection via Multi-Agent Behavioral Contract Synthesis
by: Wang, Yongchao, et al.
Published: (2026)
by: Wang, Yongchao, et al.
Published: (2026)
Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
by: Chen, Xuan, et al.
Published: (2026)
by: Chen, Xuan, et al.
Published: (2026)
RiskTagger: An LLM-based Agent for Automatic Annotation of Web3 Crypto Money Laundering Behaviors
by: Lin, Dan, et al.
Published: (2025)
by: Lin, Dan, et al.
Published: (2025)
FuzzingBrain V2: A Multi-Agent LLM System for Automated Vulnerability Discovery and Reproduction
by: Sheng, Ze, et al.
Published: (2026)
by: Sheng, Ze, et al.
Published: (2026)
QASecClaw: A Multi-Agent LLM Approach for False Positive Reduction in Static Application Security Testing
by: Ameen, Mohd Ruhul, et al.
Published: (2026)
by: Ameen, Mohd Ruhul, et al.
Published: (2026)
Pinning Is Futile: You Need More Than Local Dependency Versioning to Defend against Supply Chain Attacks
by: He, Hao, et al.
Published: (2025)
by: He, Hao, et al.
Published: (2025)
Proving and Rewarding Client Diversity to Strengthen Resilience of Blockchain Networks
by: Ron, Javier, et al.
Published: (2024)
by: Ron, Javier, et al.
Published: (2024)
RiskHarvester: A Risk-based Tool to Prioritize Secret Removal Efforts in Software Artifacts
by: Basak, Setu Kumar, et al.
Published: (2025)
by: Basak, Setu Kumar, et al.
Published: (2025)
Similar Items
-
When "Correct" Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents?
by: Peng, Yibo, et al.
Published: (2025) -
When Specifications Meet Reality: Uncovering API Inconsistencies in Ethereum Infrastructure
by: Ma, Jie, et al.
Published: (2026) -
When MCP Servers Attack: Taxonomy, Feasibility, and Mitigation
by: Zhao, Weibo, et al.
Published: (2025) -
Take a Step Further: Understanding Page Spray in Linux Kernel Exploitation
by: Guo, Ziyi, et al.
Published: (2024) -
When Security Meets Usability: An Empirical Investigation of Post-Quantum Cryptography APIs
by: Toruan, Marthin, et al.
Published: (2026)