Can Model Internals Detect MCP Tool Poisoning That Text Analysis Cannot?
Fuente:
Zenodo
Gespeichert in:
| 1. Verfasser: | |
|---|---|
| Format: | Recurso digital |
| Veröffentlicht: |
Zenodo
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866901587991986176 |
|---|---|
| author | Leung, Wan Sheng |
| author_facet | Leung, Wan Sheng |
| contents | I investigate whether looking inside a model's activations can catch poisoned MCP tool descriptions better than text scanning. On a dataset where safe and malicious descriptions cover the same topics with heavily overlapping vocabulary, text classifiers top out at 72-79%. A simple logistic regression trained on GPT-2's internal activations hits 97-98.5% and stays at 97% even after removing the effect of text length. Statistically significant (p=0.005). But this is GPT-2, not Claude, and 200 LLM-generated samples, not production data. The next step is SAE analysis on a real model. |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_19990741 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Can Model Internals Detect MCP Tool Poisoning That Text Analysis Cannot? Leung, Wan Sheng activation probes MCP security tool poisoning model internals AI safety mechanistic interpretability I investigate whether looking inside a model's activations can catch poisoned MCP tool descriptions better than text scanning. On a dataset where safe and malicious descriptions cover the same topics with heavily overlapping vocabulary, text classifiers top out at 72-79%. A simple logistic regression trained on GPT-2's internal activations hits 97-98.5% and stays at 97% even after removing the effect of text length. Statistically significant (p=0.005). But this is GPT-2, not Claude, and 200 LLM-generated samples, not production data. The next step is SAE analysis on a real model. |
| title | Can Model Internals Detect MCP Tool Poisoning That Text Analysis Cannot? |
| topic | activation probes MCP security tool poisoning model internals AI safety mechanistic interpretability |
| url | https://doi.org/10.5281/zenodo.19990741 |