Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Xuan, Yan, Lu, Zhang, Ruqi, Zhang, Xiangyu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Identifying Adversary Tactics and Techniques in Malware Binaries with an LLM Agent
by: Xuan, Zhou, et al.
Published: (2026)
by: Xuan, Zhou, et al.
Published: (2026)
PentestGPT: An LLM-empowered Automatic Penetration Testing Tool
by: Deng, Gelei, et al.
Published: (2023)
by: Deng, Gelei, et al.
Published: (2023)
Clawdrain: Exploiting Tool-Calling Chains for Stealthy Token Exhaustion in OpenClaw Agents
by: Dong, Ben, et al.
Published: (2026)
by: Dong, Ben, et al.
Published: (2026)
How Can ChatGPT Support Human Security Testers to Help Mitigate Supply Chain Attacks?
by: Zhang, Ying, et al.
Published: (2023)
by: Zhang, Ying, et al.
Published: (2023)
SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration
by: Guo, Zihan, et al.
Published: (2026)
by: Guo, Zihan, et al.
Published: (2026)
SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
by: Hu, Qi, et al.
Published: (2026)
by: Hu, Qi, et al.
Published: (2026)
What Makes a Good LLM Agent for Real-world Penetration Testing?
by: Deng, Gelei, et al.
Published: (2026)
by: Deng, Gelei, et al.
Published: (2026)
Auditing MCP Servers for Over-Privileged Tool Capabilities
by: Huang, Charoes, et al.
Published: (2026)
by: Huang, Charoes, et al.
Published: (2026)
HarnessAgent: Scaling Automatic Fuzzing Harness Construction with Tool-Augmented LLM Pipelines
by: Yang, Kang, et al.
Published: (2025)
by: Yang, Kang, et al.
Published: (2025)
RAN Tester UE: An Automated Declarative UE Centric Security Testing Platform
by: Ueltschey, Charles Marion, et al.
Published: (2025)
by: Ueltschey, Charles Marion, et al.
Published: (2025)
Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications
by: Lu, Xiaoyue, et al.
Published: (2026)
by: Lu, Xiaoyue, et al.
Published: (2026)
RACC: Representation-Aware Coverage Criteria for LLM Safety Testing
by: Wei, Zeming, et al.
Published: (2026)
by: Wei, Zeming, et al.
Published: (2026)
ChainFuzzer: Greybox Fuzzing for Workflow-Level Multi-Tool Vulnerabilities in LLM Agents
by: Wu, Jiangrong, et al.
Published: (2026)
by: Wu, Jiangrong, et al.
Published: (2026)
Coverage-Guided Multi-Agent Harness Generation for Java Library Fuzzing
by: Loose, Nils, et al.
Published: (2026)
by: Loose, Nils, et al.
Published: (2026)
Probing Privacy Leaks in LLM-based Code Generation via Test Generation
by: Ge, Yifei, et al.
Published: (2026)
by: Ge, Yifei, et al.
Published: (2026)
SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs
by: Xia, Hongfei, et al.
Published: (2025)
by: Xia, Hongfei, et al.
Published: (2025)
Usability as a Weapon: Attacking the Safety of LLM-Based Code Generation via Usability Requirements
by: Li, Yue, et al.
Published: (2026)
by: Li, Yue, et al.
Published: (2026)
A Systematic Study of LLM-Based Architectures for Automated Patching
by: Xu, Qingxiao, et al.
Published: (2026)
by: Xu, Qingxiao, et al.
Published: (2026)
Towards the Systematic Testing of Regular Expression Engines
by: Çakar, Berk, et al.
Published: (2026)
by: Çakar, Berk, et al.
Published: (2026)
QASecClaw: A Multi-Agent LLM Approach for False Positive Reduction in Static Application Security Testing
by: Ameen, Mohd Ruhul, et al.
Published: (2026)
by: Ameen, Mohd Ruhul, et al.
Published: (2026)
ARGUS: Defending LLM Agents Against Context-Aware Prompt Injection
by: Weng, Shihao, et al.
Published: (2026)
by: Weng, Shihao, et al.
Published: (2026)
Enhancing GUI Exploration Coverage of Android Apps with Deep Link-Integrated Monkey
by: Hu, Han, et al.
Published: (2024)
by: Hu, Han, et al.
Published: (2024)
Unity is Strength: Enhancing Precision in Reentrancy Vulnerability Detection of Smart Contract Analysis Tools
by: Wang, Zexu, et al.
Published: (2024)
by: Wang, Zexu, et al.
Published: (2024)
Who Audits the Auditor? Tamper-Proof Fraud Detection with Blockchain-Anchored Explainable ML
by: Wang, Zhaohui
Published: (2026)
by: Wang, Zhaohui
Published: (2026)
Beyond Imprecise Distance Metrics: Trace-Guided Directed Greybox Fuzzing via LLM-Predicted Call Stacks
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
RepoMark: A Data-Usage Auditing Framework for Code Large Language Models
by: Qu, Wenjie, et al.
Published: (2025)
by: Qu, Wenjie, et al.
Published: (2025)
Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing
by: Peng, Jiaren, et al.
Published: (2026)
by: Peng, Jiaren, et al.
Published: (2026)
TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories
by: Chen, Yen-Shan, et al.
Published: (2026)
by: Chen, Yen-Shan, et al.
Published: (2026)
"I Don't Use AI for Everything": Exploring Utility, Attitude, and Responsibility of AI-empowered Tools in Software Development
by: Pan, Shidong, et al.
Published: (2024)
by: Pan, Shidong, et al.
Published: (2024)
APT-ClaritySet: A Large-Scale, High-Fidelity Labeled Dataset for APT Malware with Alias Normalization and Graph-Based Deduplication
by: Yin, Zhenhao, et al.
Published: (2025)
by: Yin, Zhenhao, et al.
Published: (2025)
Security of Language Models for Code: A Systematic Literature Review
by: Chen, Yuchen, et al.
Published: (2024)
by: Chen, Yuchen, et al.
Published: (2024)
AutoDFBench 1.0: A Benchmarking Framework for Digital Forensic Tool Testing and Generated Code Evaluation
by: Wickramasekara, Akila, et al.
Published: (2025)
by: Wickramasekara, Akila, et al.
Published: (2025)
LLM-Powered Detection of Price Manipulation in DeFi
by: Liu, Lu, et al.
Published: (2025)
by: Liu, Lu, et al.
Published: (2025)
FuzzingBrain V2: A Multi-Agent LLM System for Automated Vulnerability Discovery and Reproduction
by: Sheng, Ze, et al.
Published: (2026)
by: Sheng, Ze, et al.
Published: (2026)
CHASE: LLM Agents for Dissecting Malicious PyPI Packages
by: Toda, Takaaki, et al.
Published: (2026)
by: Toda, Takaaki, et al.
Published: (2026)
LLMs as Firmware Experts: A Runtime-Grown Tree-of-Agents Framework
by: Zhang, Xiangrui, et al.
Published: (2025)
by: Zhang, Xiangrui, et al.
Published: (2025)
LLM Agents for Automated Web Vulnerability Reproduction: Are We There Yet?
by: Liu, Bin, et al.
Published: (2025)
by: Liu, Bin, et al.
Published: (2025)
From LLMs to Agents: A Comparative Evaluation of LLMs and LLM-based Agents in Security Patch Detection
by: Han, Junxiao, et al.
Published: (2025)
by: Han, Junxiao, et al.
Published: (2025)
Uncovering EDK2 Firmware Flaws: Insights from Code Audit Tools
by: Farahani, Mahsa, et al.
Published: (2024)
by: Farahani, Mahsa, et al.
Published: (2024)
Defending Code Language Models against Backdoor Attacks with Deceptive Cross-Entropy Loss
by: Yang, Guang, et al.
Published: (2024)
by: Yang, Guang, et al.
Published: (2024)
Similar Items
-
Identifying Adversary Tactics and Techniques in Malware Binaries with an LLM Agent
by: Xuan, Zhou, et al.
Published: (2026) -
PentestGPT: An LLM-empowered Automatic Penetration Testing Tool
by: Deng, Gelei, et al.
Published: (2023) -
Clawdrain: Exploiting Tool-Calling Chains for Stealthy Token Exhaustion in OpenClaw Agents
by: Dong, Ben, et al.
Published: (2026) -
How Can ChatGPT Support Human Security Testers to Help Mitigate Supply Chain Attacks?
by: Zhang, Ying, et al.
Published: (2023) -
SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration
by: Guo, Zihan, et al.
Published: (2026)