MCIP: Protecting MCP Safety via Model Contextual Integrity Protocol

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jing, Huihao, Li, Haoran, Hu, Wenbin, Hu, Qi, Xu, Heli, Chu, Tianshu, Hu, Peizhao, Song, Yangqiu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914063515123712
author Jing, Huihao
Li, Haoran
Hu, Wenbin
Hu, Qi
Xu, Heli
Chu, Tianshu
Hu, Peizhao
Song, Yangqiu
author_facet Jing, Huihao
Li, Haoran
Hu, Wenbin
Hu, Qi
Xu, Heli
Chu, Tianshu
Hu, Peizhao
Song, Yangqiu
contents As Model Context Protocol (MCP) introduces an easy-to-use ecosystem for users and developers, it also brings underexplored safety risks. Its decentralized architecture, which separates clients and servers, poses unique challenges for systematic safety analysis. This paper proposes a novel framework to enhance MCP safety. Guided by the MAESTRO framework, we first analyze the missing safety mechanisms in MCP, and based on this analysis, we propose the Model Contextual Integrity Protocol (MCIP), a refined version of MCP that addresses these gaps. Next, we develop a fine-grained taxonomy that captures a diverse range of unsafe behaviors observed in MCP scenarios. Building on this taxonomy, we develop benchmark and training data that support the evaluation and improvement of LLMs' capabilities in identifying safety risks within MCP interactions. Leveraging the proposed benchmark and training data, we conduct extensive experiments on state-of-the-art LLMs. The results highlight LLMs' vulnerabilities in MCP interactions and demonstrate that our approach substantially improves their safety performance.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14590
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MCIP: Protecting MCP Safety via Model Contextual Integrity Protocol
Jing, Huihao
Li, Haoran
Hu, Wenbin
Hu, Qi
Xu, Heli
Chu, Tianshu
Hu, Peizhao
Song, Yangqiu
Computation and Language
As Model Context Protocol (MCP) introduces an easy-to-use ecosystem for users and developers, it also brings underexplored safety risks. Its decentralized architecture, which separates clients and servers, poses unique challenges for systematic safety analysis. This paper proposes a novel framework to enhance MCP safety. Guided by the MAESTRO framework, we first analyze the missing safety mechanisms in MCP, and based on this analysis, we propose the Model Contextual Integrity Protocol (MCIP), a refined version of MCP that addresses these gaps. Next, we develop a fine-grained taxonomy that captures a diverse range of unsafe behaviors observed in MCP scenarios. Building on this taxonomy, we develop benchmark and training data that support the evaluation and improvement of LLMs' capabilities in identifying safety risks within MCP interactions. Leveraging the proposed benchmark and training data, we conduct extensive experiments on state-of-the-art LLMs. The results highlight LLMs' vulnerabilities in MCP interactions and demonstrate that our approach substantially improves their safety performance.
title MCIP: Protecting MCP Safety via Model Contextual Integrity Protocol
topic Computation and Language
url https://arxiv.org/abs/2505.14590