Automatic Red Teaming LLM-based Agents with Model Context Protocol Tools
Fuente:
arXiv
Saved in:
| Main Authors: | He, Ping, Li, Changjiang, Zhao, Binbin, Du, Tianyu, Ji, Shouling |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Security Risks of ML-based Malware Detection Systems: A Survey
by: He, Ping, et al.
Published: (2025)
by: He, Ping, et al.
Published: (2025)
Defending against Adversarial Malware Attacks on ML-based Android Malware Detection Systems
by: He, Ping, et al.
Published: (2025)
by: He, Ping, et al.
Published: (2025)
MCPXKIT: The Unified Toolkit for Analyzing Model Context Protocol Security
by: Guo, Yongjian, et al.
Published: (2025)
by: Guo, Yongjian, et al.
Published: (2025)
TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs
by: Shen, Qingchao, et al.
Published: (2026)
by: Shen, Qingchao, et al.
Published: (2026)
Chimera: Harnessing Multi-Agent LLMs for Automatic Insider Threat Simulation
by: Yu, Jiongchi, et al.
Published: (2025)
by: Yu, Jiongchi, et al.
Published: (2025)
Automatically Generating Rules of Malicious Software Packages via Large Language Model
by: Zhang, XiangRui, et al.
Published: (2025)
by: Zhang, XiangRui, et al.
Published: (2025)
SCRUTINEER: Detecting Logic-Level Usage Violations of Reusable Components in Smart Contracts
by: Lin, Xingshuang, et al.
Published: (2025)
by: Lin, Xingshuang, et al.
Published: (2025)
VULPO: Context-Aware Vulnerability Detection via On-Policy LLM Optimization
by: Li, Youpeng, et al.
Published: (2025)
by: Li, Youpeng, et al.
Published: (2025)
Beyond Classification: Evaluating LLMs for Fine-Grained Automatic Malware Behavior Auditing
by: Zheng, Xinran, et al.
Published: (2025)
by: Zheng, Xinran, et al.
Published: (2025)
SecureFixAgent: A Hybrid LLM Agent for Automated Python Static Vulnerability Repair
by: Gajjar, Jugal, et al.
Published: (2025)
by: Gajjar, Jugal, et al.
Published: (2025)
Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges
by: Al-Kaswan, Ali, et al.
Published: (2026)
by: Al-Kaswan, Ali, et al.
Published: (2026)
PROMFUZZ: Leveraging LLM-Driven and Bug-Oriented Composite Analysis for Detecting Functional Bugs in Smart Contracts
by: Lin, Xingshuang, et al.
Published: (2025)
by: Lin, Xingshuang, et al.
Published: (2025)
LLM-Agent-UMF: LLM-based Agent Unified Modeling Framework for Seamless Design of Multi Active/Passive Core-Agent Architectures
by: Hassouna, Amine Ben, et al.
Published: (2024)
by: Hassouna, Amine Ben, et al.
Published: (2024)
Data and Context Matter: Towards Generalizing AI-based Software Vulnerability Detection
by: Safdar, Rijha, et al.
Published: (2025)
by: Safdar, Rijha, et al.
Published: (2025)
Taint-Style Vulnerability Detection and Confirmation for Node.js Packages Using LLM Agent Reasoning
by: Ni, Ronghao, et al.
Published: (2026)
by: Ni, Ronghao, et al.
Published: (2026)
Zero-Permission Manipulation: Can We Trust Large Multimodal Model Powered GUI Agents?
by: Qian, Yi, et al.
Published: (2026)
by: Qian, Yi, et al.
Published: (2026)
Decaf: Improving Neural Decompilation with Automatic Feedback and Search
by: Shypula, Alexander, et al.
Published: (2026)
by: Shypula, Alexander, et al.
Published: (2026)
Unveiling the Landscape of LLM Deployment in the Wild: An Empirical Study
by: Hou, Xinyi, et al.
Published: (2025)
by: Hou, Xinyi, et al.
Published: (2025)
HarnessAgent: Scaling Automatic Fuzzing Harness Construction with Tool-Augmented LLM Pipelines
by: Yang, Kang, et al.
Published: (2025)
by: Yang, Kang, et al.
Published: (2025)
Static Semantics Reconstruction for Enhancing JavaScript-WebAssembly Multilingual Malware Detection
by: Xia, Yifan, et al.
Published: (2023)
by: Xia, Yifan, et al.
Published: (2023)
QLPro: Automated Code Vulnerability Discovery via LLM and Static Code Analysis Integration
by: Hu, Junze, et al.
Published: (2025)
by: Hu, Junze, et al.
Published: (2025)
Model Context Protocol Threat Modeling and Analyzing Vulnerabilities to Prompt Injection with Tool Poisoning
by: Huang, Charoes, et al.
Published: (2026)
by: Huang, Charoes, et al.
Published: (2026)
PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design
by: Yang, Ruozhao, et al.
Published: (2025)
by: Yang, Ruozhao, et al.
Published: (2025)
From SFT to RL: Demystifying the Post-Training Pipeline for LLM-based Vulnerability Detection
by: Li, Youpeng, et al.
Published: (2026)
by: Li, Youpeng, et al.
Published: (2026)
OpenSage: Self-programming Agent Generation Engine
by: Li, Hongwei, et al.
Published: (2026)
by: Li, Hongwei, et al.
Published: (2026)
VulDetectBench: Evaluating the Deep Capability of Vulnerability Detection with Large Language Models
by: Liu, Yu, et al.
Published: (2024)
by: Liu, Yu, et al.
Published: (2024)
Comment and Control: Hijacking Agentic Workflows via Context-Grounded Evolution
by: Fendley, Neil, et al.
Published: (2026)
by: Fendley, Neil, et al.
Published: (2026)
LLAMAFUZZ: Large Language Model Enhanced Greybox Fuzzing
by: Zhang, Hongxiang, et al.
Published: (2024)
by: Zhang, Hongxiang, et al.
Published: (2024)
An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong Detection
by: Yan, Shenao, et al.
Published: (2024)
by: Yan, Shenao, et al.
Published: (2024)
Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing
by: Peng, Jiaren, et al.
Published: (2026)
by: Peng, Jiaren, et al.
Published: (2026)
Implicit Patterns in LLM-Based Binary Analysis
by: Li, Qiang, et al.
Published: (2026)
by: Li, Qiang, et al.
Published: (2026)
PentestGPT: An LLM-empowered Automatic Penetration Testing Tool
by: Deng, Gelei, et al.
Published: (2023)
by: Deng, Gelei, et al.
Published: (2023)
Reflection-Driven Control for Trustworthy Code Agents
by: Wang, Bin, et al.
Published: (2025)
by: Wang, Bin, et al.
Published: (2025)
KernelGPT: Enhanced Kernel Fuzzing via Large Language Models
by: Yang, Chenyuan, et al.
Published: (2023)
by: Yang, Chenyuan, et al.
Published: (2023)
Security of LLM-generated Code: A Comparative Analysis
by: Morkonda, Srivathsan G, et al.
Published: (2026)
by: Morkonda, Srivathsan G, et al.
Published: (2026)
The potential of LLM-generated reports in DevSecOps
by: Lykousas, Nikolaos, et al.
Published: (2024)
by: Lykousas, Nikolaos, et al.
Published: (2024)
Favia: Forensic Agent for Vulnerability-fix Identification and Analysis
by: Storhaug, André, et al.
Published: (2026)
by: Storhaug, André, et al.
Published: (2026)
SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers
by: Qin, Kaihua, et al.
Published: (2026)
by: Qin, Kaihua, et al.
Published: (2026)
LLM-enabled Applications Require System-Level Threat Monitoring
by: Zhang, Yedi, et al.
Published: (2026)
by: Zhang, Yedi, et al.
Published: (2026)
SKILLS: Structured Knowledge Injection for LLM-Driven Telecommunications Operations
by: Brett, Ivo
Published: (2026)
by: Brett, Ivo
Published: (2026)
Similar Items
-
On the Security Risks of ML-based Malware Detection Systems: A Survey
by: He, Ping, et al.
Published: (2025) -
Defending against Adversarial Malware Attacks on ML-based Android Malware Detection Systems
by: He, Ping, et al.
Published: (2025) -
MCPXKIT: The Unified Toolkit for Analyzing Model Context Protocol Security
by: Guo, Yongjian, et al.
Published: (2025) -
TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs
by: Shen, Qingchao, et al.
Published: (2026) -
Chimera: Harnessing Multi-Agent LLMs for Automatic Insider Threat Simulation
by: Yu, Jiongchi, et al.
Published: (2025)