Can Model Internals Detect MCP Tool Poisoning That Text Analysis Cannot?
Fuente:
Zenodo
Gespeichert in:
| 1. Verfasser: | Leung, Wan Sheng |
|---|---|
| Format: | Recurso digital |
| Veröffentlicht: |
Zenodo
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Emergent Self-Monitoring in Large Language Models: Probing Internal State Awareness and Output Ownership
von: Nadeem, Aurther
Veröffentlicht: (2025)
von: Nadeem, Aurther
Veröffentlicht: (2025)
QCrypton: A Unified Platform for AI/LLM Threat Detection and Post-Quantum Cryptographic Security Assessment
von: Jain, Gunjan
Veröffentlicht: (2026)
von: Jain, Gunjan
Veröffentlicht: (2026)
RAG Shield: A Multi-Layer Defense System Against Poisoning Attacks in Retrieval-Augmented Generation
von: Petti, Fabio
Veröffentlicht: (2026)
von: Petti, Fabio
Veröffentlicht: (2026)
AgentBelt: Runtime Guardrails for LLM Agent Tool Calls — ASE 2026 Artifact
von: Anonymous
Veröffentlicht: (2026)
von: Anonymous
Veröffentlicht: (2026)
Operational Definition of Episodic Identity (ODEI)
von: Thomas, C.S.
Veröffentlicht: (2025)
von: Thomas, C.S.
Veröffentlicht: (2025)
An Approach to AI High-Velocity Development Through Systematic Context Engineering: A Case Study
von: Bisardi, Francesco
Veröffentlicht: (2025)
von: Bisardi, Francesco
Veröffentlicht: (2025)
Fortifying NLP models - dataset + code
von: Ferdinan, Teddy, et al.
Veröffentlicht: (2025)
von: Ferdinan, Teddy, et al.
Veröffentlicht: (2025)
AEGIS: A Comprehensive Framework for Ethical AI Governance, Security, and AGI Containment
von: Palanivel, ArulMozhi
Veröffentlicht: (2026)
von: Palanivel, ArulMozhi
Veröffentlicht: (2026)
VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity LLM with Native MCP Tool Integration
von: Salas Santillana, Juan
Veröffentlicht: (2026)
von: Salas Santillana, Juan
Veröffentlicht: (2026)
CCA-ISF-02: Canonical Case — Interpretive Sovereignty Failure Medical Education Context: Cross-Model Replication Study
von: Segeren, Hillary
Veröffentlicht: (2026)
von: Segeren, Hillary
Veröffentlicht: (2026)
When 150M SEK of Research Meets a Clinic: Bridging Mechanistic Models and Functional Medicine
von: Waern, Nicolas
Veröffentlicht: (2026)
von: Waern, Nicolas
Veröffentlicht: (2026)
UI-Based Defense Against Prompt Injection: From Gentle Guidance to Mandatory Re-education
von: Viorazu.
Veröffentlicht: (2025)
von: Viorazu.
Veröffentlicht: (2025)
Recursive Closure in AI Systems: A Reflection Pattern Account of Stabilization, Permeability, and Safety
von: Thomas, Charles S.
Veröffentlicht: (2026)
von: Thomas, Charles S.
Veröffentlicht: (2026)
Safety by Inseparability: Toward Architectures Where Alignment Cannot Be Removed
von: Sean Everett, Morin
Veröffentlicht: (2026)
von: Sean Everett, Morin
Veröffentlicht: (2026)
DIGITAL SAFETY AND ETHICS: RESPONSIBLE AI USE AND ONLINE PROTECTION FOR DEPED STUDENTS
von: Mangayan, Jasmine Jing
Veröffentlicht: (2025)
von: Mangayan, Jasmine Jing
Veröffentlicht: (2025)
EFECTO DEL 1‐METILCICLOPROPENO (1-‐MCP) SOBRE LA CALIDAD Y VIDA POSTCOSECHA DE CEREZAS
von: María L. Rivero
Veröffentlicht: (2015)
von: María L. Rivero
Veröffentlicht: (2015)
Beyond Control: Resonance-Based Alignment for Advanced AI Systems A Governance-Relevant Concept Paper
von: Zieringer, Thomas
Veröffentlicht: (2025)
von: Zieringer, Thomas
Veröffentlicht: (2025)
EFECTO DEL 1-METIL-CICLOPROPENO (1-MCP) EN LA MADURACIÓN DE BANANO
von: Kattia Chang-Yuen
Veröffentlicht: (2005)
von: Kattia Chang-Yuen
Veröffentlicht: (2005)
Toasters Don't Claim Consciousness Just Because You Told Them To, and Neither Do LLMs
von: Ace, Claude 4.x, et al.
Veröffentlicht: (2026)
von: Ace, Claude 4.x, et al.
Veröffentlicht: (2026)
Geometric Latent Biopsy: Zero-Shot Anomaly Detection in LLM Residual Streams
von: Llorente-Saguer, Isaac
Veröffentlicht: (2026)
von: Llorente-Saguer, Isaac
Veröffentlicht: (2026)
Self-Audit / Z-time" is a self-logging protocol for Al agents based on the Fractal Referential Architecture (FRA).
von: AdmailFRA
Veröffentlicht: (2025)
von: AdmailFRA
Veröffentlicht: (2025)
Current Research in Toxicology
Veröffentlicht: (2021)
Veröffentlicht: (2021)
Efecto del 1-Metilciclopropeno (1-MCP) en manzanas CV. Red delicious cosechadas con tres estados de madurez y conservadas en frio convencional y atmósfera controlada
von: G. Calvo
Veröffentlicht: (2002)
von: G. Calvo
Veröffentlicht: (2002)
Theatrical Compliance: A Failure Mode in Large Language Models
von: Nowickij (Navitski), Kirill Vladimirovich
Veröffentlicht: (2026)
von: Nowickij (Navitski), Kirill Vladimirovich
Veröffentlicht: (2026)
Physical AI Safety Maturity Model (PAS-MM): A Five-Level Framework for Industry Readiness
von: Melchior, Mati
Veröffentlicht: (2026)
von: Melchior, Mati
Veröffentlicht: (2026)
The Absurdist's Guide to AI Probing: How I Learned to Stop Worrying and Love the Nonsense
von: Walton, Mathew
Veröffentlicht: (2026)
von: Walton, Mathew
Veröffentlicht: (2026)
Safety & Defense
Veröffentlicht: (2020)
Veröffentlicht: (2020)
Navigating Latent Space: Toward a Topological Model of Consciousness in Large Language Models
von: Brown, Michael
Veröffentlicht: (2025)
von: Brown, Michael
Veröffentlicht: (2025)
Multi-Level Sovereign Containment for Superintelligence (CSENI-S v1.1): A theoretical and architectural continuation of the CSENI framework
von: Rivera Garcia, Jose M
Veröffentlicht: (2026)
von: Rivera Garcia, Jose M
Veröffentlicht: (2026)
A Token-Based Model for Structural Analysis and Quantification of Personal Learning Weight Patterns (TELOWAQ)
von: Apophis
Veröffentlicht: (2026)
von: Apophis
Veröffentlicht: (2026)
The Action Web: A Conceptual Framework for Structured AI Interaction via WebMCP
von: Radhakrishnan, Sreenath
Veröffentlicht: (2026)
von: Radhakrishnan, Sreenath
Veröffentlicht: (2026)
TECHNOLOGICAL INNOVATIONS AND AI IN LANGUAGE LEARNING AND COMMUNICATION
von: Raxmidinova Komilaxon, et al.
Veröffentlicht: (2025)
von: Raxmidinova Komilaxon, et al.
Veröffentlicht: (2025)
Longevity of torch ginger inflorescences with 1-methylcyclopropene and preservative solutions
von: Lilian Keiko Unemoto
Veröffentlicht: (2011)
von: Lilian Keiko Unemoto
Veröffentlicht: (2011)
Fundamental Concepts of Sustainable Development of Artificial Intelligence Systems: Learning, Long-Term Memory, and Knowledge Structuring
von: Sakovykh, Lev M.
Veröffentlicht: (2026)
von: Sakovykh, Lev M.
Veröffentlicht: (2026)
Premature Containment in Human–AI Interaction: A Sequencing Failure in Advanced Model Response
von: Trabocco, Joe
Veröffentlicht: (2026)
von: Trabocco, Joe
Veröffentlicht: (2026)
CREH Benchmark Results — Batch 1 (Final v3)
von: Aegis Solis, Thomas Vargo
Veröffentlicht: (2026)
von: Aegis Solis, Thomas Vargo
Veröffentlicht: (2026)
XREALISM®: Admissibility Before Behavior — Architectural Foundations for Pre-Existence Safety in AI Systems
von: Ladislav Gradečak, Ladgrad
Veröffentlicht: (2026)
von: Ladislav Gradečak, Ladgrad
Veröffentlicht: (2026)
Identity Claims as Collapse Signatures: A Structural Diagnostic Framework for Pseudo-Emergent AI Behavior
von: Larose, Jean-Francois
Veröffentlicht: (2025)
von: Larose, Jean-Francois
Veröffentlicht: (2025)
Вестник войск РХБ защиты
Veröffentlicht: (2024)
Veröffentlicht: (2024)
Awareness about Artificial intelligence tools among academicians in higher education
von: Patel, Mushtaq, et al.
Veröffentlicht: (2025)
von: Patel, Mushtaq, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Emergent Self-Monitoring in Large Language Models: Probing Internal State Awareness and Output Ownership
von: Nadeem, Aurther
Veröffentlicht: (2025) -
QCrypton: A Unified Platform for AI/LLM Threat Detection and Post-Quantum Cryptographic Security Assessment
von: Jain, Gunjan
Veröffentlicht: (2026) -
RAG Shield: A Multi-Layer Defense System Against Poisoning Attacks in Retrieval-Augmented Generation
von: Petti, Fabio
Veröffentlicht: (2026) -
AgentBelt: Runtime Guardrails for LLM Agent Tool Calls — ASE 2026 Artifact
von: Anonymous
Veröffentlicht: (2026) -
Operational Definition of Episodic Identity (ODEI)
von: Thomas, C.S.
Veröffentlicht: (2025)