Can Model Internals Detect MCP Tool Poisoning That Text Analysis Cannot?

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Leung, Wan Sheng
Format: Recurso digital
Veröffentlicht: Zenodo 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866901587991986176
author Leung, Wan Sheng
author_facet Leung, Wan Sheng
contents I investigate whether looking inside a model's activations can catch poisoned MCP tool descriptions better than text scanning. On a dataset where safe and malicious descriptions cover the same topics with heavily overlapping vocabulary, text classifiers top out at 72-79%. A simple logistic regression trained on GPT-2's internal activations hits 97-98.5% and stays at 97% even after removing the effect of text length. Statistically significant (p=0.005). But this is GPT-2, not Claude, and 200 LLM-generated samples, not production data. The next step is SAE analysis on a real model.
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_19990741
institution Zenodo
language
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Can Model Internals Detect MCP Tool Poisoning That Text Analysis Cannot?
Leung, Wan Sheng
activation probes
MCP security
tool poisoning
model internals
AI safety
mechanistic interpretability
I investigate whether looking inside a model's activations can catch poisoned MCP tool descriptions better than text scanning. On a dataset where safe and malicious descriptions cover the same topics with heavily overlapping vocabulary, text classifiers top out at 72-79%. A simple logistic regression trained on GPT-2's internal activations hits 97-98.5% and stays at 97% even after removing the effect of text length. Statistically significant (p=0.005). But this is GPT-2, not Claude, and 200 LLM-generated samples, not production data. The next step is SAE analysis on a real model.
title Can Model Internals Detect MCP Tool Poisoning That Text Analysis Cannot?
topic activation probes
MCP security
tool poisoning
model internals
AI safety
mechanistic interpretability
url https://doi.org/10.5281/zenodo.19990741