MCP Safety Audit: LLMs with the Model Context Protocol Allow Major Security Exploits
Fuente:
arXiv
Saved in:
| Main Authors: | Radosevich, Brandon, Halloran, John |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment
by: Halloran, John
Published: (2025)
by: Halloran, John
Published: (2025)
Understanding the Effects of Safety Unalignment on Large Language Models
by: Halloran, John T.
Published: (2026)
by: Halloran, John T.
Published: (2026)
Leveraging RAG for Training-Free Alignment of LLMs
by: Halloran, John T.
Published: (2026)
by: Halloran, John T.
Published: (2026)
Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks
by: Halloran, John T., et al.
Published: (2026)
by: Halloran, John T., et al.
Published: (2026)
Enterprise-Grade Security for the Model Context Protocol (MCP): Frameworks and Mitigation Strategies
by: Narajala, Vineeth Sai, et al.
Published: (2025)
by: Narajala, Vineeth Sai, et al.
Published: (2025)
Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions
by: Hou, Xinyi, et al.
Published: (2025)
by: Hou, Xinyi, et al.
Published: (2025)
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
by: Zhang, Dongsen, et al.
Published: (2025)
by: Zhang, Dongsen, et al.
Published: (2025)
MCP-DPT: A Defense-Placement Taxonomy and Coverage Analysis for Model Context Protocol Security
by: Rostamzadeh, Mehrdad, et al.
Published: (2026)
by: Rostamzadeh, Mehrdad, et al.
Published: (2026)
SAMEP: A Secure Protocol for Persistent Context Sharing Across AI Agents
by: Masoor, Hari
Published: (2025)
by: Masoor, Hari
Published: (2025)
Systematization of Knowledge: Security and Safety in the Model Context Protocol Ecosystem
by: Gaire, Shiva, et al.
Published: (2025)
by: Gaire, Shiva, et al.
Published: (2025)
MCP-Guard: A Multi-Stage Defense-in-Depth Framework for Securing Model Context Protocol in Agentic AI
by: Xing, Wenpeng, et al.
Published: (2025)
by: Xing, Wenpeng, et al.
Published: (2025)
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
by: Wang, Zhun, et al.
Published: (2026)
by: Wang, Zhun, et al.
Published: (2026)
Privacy Auditing of Large Language Models
by: Panda, Ashwinee, et al.
Published: (2025)
by: Panda, Ashwinee, et al.
Published: (2025)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
by: Jiang, Yukun, et al.
Published: (2026)
by: Jiang, Yukun, et al.
Published: (2026)
Sovereign Context Protocol: An Open Attribution Layer for Human-Generated Content in the Age of Large Language Models
by: Panchigar, Praneel, et al.
Published: (2026)
by: Panchigar, Praneel, et al.
Published: (2026)
A Safety and Security Framework for Real-World Agentic Systems
by: Ghosh, Shaona, et al.
Published: (2025)
by: Ghosh, Shaona, et al.
Published: (2025)
Secure Multi-Modal Data Fusion in Federated Digital Health Systems via MCP
by: Aueawatthanaphisut, Aueaphum
Published: (2025)
by: Aueawatthanaphisut, Aueaphum
Published: (2025)
Exploiting LLM Quantization
by: Egashira, Kazuki, et al.
Published: (2024)
by: Egashira, Kazuki, et al.
Published: (2024)
Fast Exact Unlearning for In-Context Learning Data for LLMs
by: Muresanu, Andrei I., et al.
Published: (2024)
by: Muresanu, Andrei I., et al.
Published: (2024)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
MCP Bridge: A Lightweight, LLM-Agnostic RESTful Proxy for Model Context Protocol Servers
by: Ahmadi, Arash, et al.
Published: (2025)
by: Ahmadi, Arash, et al.
Published: (2025)
LAMD: Context-driven Android Malware Detection and Classification with LLMs
by: Qian, Xingzhi, et al.
Published: (2025)
by: Qian, Xingzhi, et al.
Published: (2025)
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
by: Panfilov, Alexander, et al.
Published: (2025)
by: Panfilov, Alexander, et al.
Published: (2025)
Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
by: Bazinska, Julia, et al.
Published: (2025)
by: Bazinska, Julia, et al.
Published: (2025)
TracLLM: A Generic Framework for Attributing Long Context LLMs
by: Wang, Yanting, et al.
Published: (2025)
by: Wang, Yanting, et al.
Published: (2025)
MCP-38: A Comprehensive Threat Taxonomy for Model Context Protocol Systems (v1.0)
by: Shen, Yi Ting, et al.
Published: (2026)
by: Shen, Yi Ting, et al.
Published: (2026)
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
by: Lu, Ning, et al.
Published: (2025)
by: Lu, Ning, et al.
Published: (2025)
Tight Auditing of Differential Privacy in MST and AIM
by: Ganev, Georgi, et al.
Published: (2026)
by: Ganev, Georgi, et al.
Published: (2026)
SME-TEAM: Leveraging Trust and Ethics for Secure and Responsible Use of AI and LLMs in SMEs
by: Sarker, Iqbal H., et al.
Published: (2025)
by: Sarker, Iqbal H., et al.
Published: (2025)
Privacy Auditing of Multi-domain Graph Pre-trained Model under Membership Inference Attacks
by: Luo, Jiayi, et al.
Published: (2025)
by: Luo, Jiayi, et al.
Published: (2025)
SafeRedirect: Defeating Internal Safety Collapse via Task-Completion Redirection in Frontier LLMs
by: Pan, Chao, et al.
Published: (2026)
by: Pan, Chao, et al.
Published: (2026)
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
Who's the Evil Twin? Differential Auditing for Undesired Behavior
by: Balappanawar, Ishwar, et al.
Published: (2025)
by: Balappanawar, Ishwar, et al.
Published: (2025)
Exploiting Efficiency Vulnerabilities in Dynamic Deep Learning Systems
by: Rathnasuriya, Ravishka, et al.
Published: (2025)
by: Rathnasuriya, Ravishka, et al.
Published: (2025)
Hound: Relation-First Knowledge Graphs for Complex-System Reasoning in Security Audits
by: Mueller, Bernhard
Published: (2025)
by: Mueller, Bernhard
Published: (2025)
SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification
by: Lai, Zhenglin, et al.
Published: (2025)
by: Lai, Zhenglin, et al.
Published: (2025)
Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
by: Huang, Tiansheng, et al.
Published: (2025)
by: Huang, Tiansheng, et al.
Published: (2025)
An In-Depth Analysis of Cyber Attacks in Secured Platforms
by: Ozoh, Parick, et al.
Published: (2025)
by: Ozoh, Parick, et al.
Published: (2025)
Exploiting Layer-Specific Vulnerabilities to Backdoor Attack in Federated Learning
by: Foroughi, Mohammad Hadi, et al.
Published: (2026)
by: Foroughi, Mohammad Hadi, et al.
Published: (2026)
Semantic Attacks on Tool-Augmented LLMs: Securing the Model Context Protocol Against Descriptor-Level Manipulation
by: Jamshidi, Saeid, et al.
Published: (2025)
by: Jamshidi, Saeid, et al.
Published: (2025)
Similar Items
-
MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment
by: Halloran, John
Published: (2025) -
Understanding the Effects of Safety Unalignment on Large Language Models
by: Halloran, John T.
Published: (2026) -
Leveraging RAG for Training-Free Alignment of LLMs
by: Halloran, John T.
Published: (2026) -
Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks
by: Halloran, John T., et al.
Published: (2026) -
Enterprise-Grade Security for the Model Context Protocol (MCP): Frameworks and Mitigation Strategies
by: Narajala, Vineeth Sai, et al.
Published: (2025)