Breaking MCP with Function Hijacking Attacks: Novel Threats for Function Calling and Agentic Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Belkhiter, Yannis, Zizzo, Giulio, Maffeis, Sergio, Tirupathi, Seshu, Kelleher, John D. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HarmLevelBench: Evaluating Harm-Level Compliance and the Impact of Quantization on Model Alignment
von: Belkhiter, Yannis, et al.
Veröffentlicht: (2024)
von: Belkhiter, Yannis, et al.
Veröffentlicht: (2024)
Step-Tagging: Toward controlling the generation of Language Reasoning Models through step monitoring
von: Belkhiter, Yannis, et al.
Veröffentlicht: (2025)
von: Belkhiter, Yannis, et al.
Veröffentlicht: (2025)
Blue Teaming Function-Calling Agents
von: Dolcetti, Greta, et al.
Veröffentlicht: (2026)
von: Dolcetti, Greta, et al.
Veröffentlicht: (2026)
TRACES: Tagging Reasoning Steps for Adaptive Cost-Efficient Early-Stopping
von: Belkhiter, Yannis, et al.
Veröffentlicht: (2026)
von: Belkhiter, Yannis, et al.
Veröffentlicht: (2026)
Pre-Hoc Predictions in AutoML: Leveraging LLMs to Enhance Model Selection and Benchmarking for Tabular datasets
von: Belkhiter, Yannis, et al.
Veröffentlicht: (2025)
von: Belkhiter, Yannis, et al.
Veröffentlicht: (2025)
Elevating Defenses: Bridging Adversarial Training and Watermarking for Model Resilience
von: Thakkar, Janvi, et al.
Veröffentlicht: (2023)
von: Thakkar, Janvi, et al.
Veröffentlicht: (2023)
Dynamic Features Adaptation in Networking: Toward Flexible training and Explainable inference
von: Belkhiter, Yannis, et al.
Veröffentlicht: (2025)
von: Belkhiter, Yannis, et al.
Veröffentlicht: (2025)
Towards a Practical Defense against Adversarial Attacks on Deep Learning-based Malware Detectors via Randomized Smoothing
von: Gibert, Daniel, et al.
Veröffentlicht: (2023)
von: Gibert, Daniel, et al.
Veröffentlicht: (2023)
A Robust Defense against Adversarial Attacks on Deep Learning-based Malware Detectors via (De)Randomized Smoothing
von: Gibert, Daniel, et al.
Veröffentlicht: (2024)
von: Gibert, Daniel, et al.
Veröffentlicht: (2024)
A Survey on Agentic Security: Applications, Threats and Defenses
von: Shahriar, Asif, et al.
Veröffentlicht: (2025)
von: Shahriar, Asif, et al.
Veröffentlicht: (2025)
Make Split, not Hijack: Preventing Feature-Space Hijacking Attacks in Split Learning
von: Khan, Tanveer, et al.
Veröffentlicht: (2024)
von: Khan, Tanveer, et al.
Veröffentlicht: (2024)
Beyond the Tip of Efficiency: Uncovering the Submerged Threats of Jailbreak Attacks in Small Language Models
von: Yi, Sibo, et al.
Veröffentlicht: (2025)
von: Yi, Sibo, et al.
Veröffentlicht: (2025)
PARASITE: Conditional System Prompt Poisoning to Hijack LLMs
von: Pham, Viet, et al.
Veröffentlicht: (2025)
von: Pham, Viet, et al.
Veröffentlicht: (2025)
In-Context Representation Hijacking
von: Yona, Itay, et al.
Veröffentlicht: (2025)
von: Yona, Itay, et al.
Veröffentlicht: (2025)
HijackRAG: Hijacking Attacks against Retrieval-Augmented Large Language Models
von: Zhang, Yucheng, et al.
Veröffentlicht: (2024)
von: Zhang, Yucheng, et al.
Veröffentlicht: (2024)
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
von: Chae, Kyubyung, et al.
Veröffentlicht: (2025)
von: Chae, Kyubyung, et al.
Veröffentlicht: (2025)
Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions
von: Hou, Xinyi, et al.
Veröffentlicht: (2025)
von: Hou, Xinyi, et al.
Veröffentlicht: (2025)
Hallucinating AI Hijacking Attack: Large Language Models and Malicious Code Recommenders
von: Noever, David, et al.
Veröffentlicht: (2024)
von: Noever, David, et al.
Veröffentlicht: (2024)
Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space
von: Huang, Yao, et al.
Veröffentlicht: (2025)
von: Huang, Yao, et al.
Veröffentlicht: (2025)
PentestMCP: A Toolkit for Agentic Penetration Testing
von: Ezetta, Zachary, et al.
Veröffentlicht: (2025)
von: Ezetta, Zachary, et al.
Veröffentlicht: (2025)
JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
von: Luo, Weidi, et al.
Veröffentlicht: (2024)
von: Luo, Weidi, et al.
Veröffentlicht: (2024)
Assessing the Impact of Packing on Machine Learning-Based Malware Detection and Classification Systems
von: Gibert, Daniel, et al.
Veröffentlicht: (2024)
von: Gibert, Daniel, et al.
Veröffentlicht: (2024)
Breaking PEFT Limitations: Leveraging Weak-to-Strong Knowledge Transfer for Backdoor Attacks in LLMs
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
Resource Consumption Threats in Large Language Models
von: Zhang, Yuanhe, et al.
Veröffentlicht: (2026)
von: Zhang, Yuanhe, et al.
Veröffentlicht: (2026)
MoJE: Mixture of Jailbreak Experts, Naive Tabular Classifiers as Guard for Prompt Attacks
von: Cornacchia, Giandomenico, et al.
Veröffentlicht: (2024)
von: Cornacchia, Giandomenico, et al.
Veröffentlicht: (2024)
Prompt-in-Content Attacks: Exploiting Uploaded Inputs to Hijack LLM Behavior
von: Lian, Zhuotao, et al.
Veröffentlicht: (2025)
von: Lian, Zhuotao, et al.
Veröffentlicht: (2025)
Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction
von: Wang, Hongtao, et al.
Veröffentlicht: (2026)
von: Wang, Hongtao, et al.
Veröffentlicht: (2026)
MCP-38: A Comprehensive Threat Taxonomy for Model Context Protocol Systems (v1.0)
von: Shen, Yi Ting, et al.
Veröffentlicht: (2026)
von: Shen, Yi Ting, et al.
Veröffentlicht: (2026)
Towards Assuring EU AI Act Compliance and Adversarial Robustness of LLMs
von: Momcilovic, Tomas Bueno, et al.
Veröffentlicht: (2024)
von: Momcilovic, Tomas Bueno, et al.
Veröffentlicht: (2024)
ThreatGPT: An Agentic AI Framework for Enhancing Public Safety through Threat Modeling
von: Zisad, Sharif Noor, et al.
Veröffentlicht: (2025)
von: Zisad, Sharif Noor, et al.
Veröffentlicht: (2025)
Blind PRNG Hijacking: An Undetectable Integrity-Preserving Attack Against LLM Watermarking
von: You, Ziyang, et al.
Veröffentlicht: (2026)
von: You, Ziyang, et al.
Veröffentlicht: (2026)
A Formal Security Framework for MCP-Based AI Agents: Threat Taxonomy, Verification Models, and Defense Mechanisms
von: Acharya, Nirajan, et al.
Veröffentlicht: (2026)
von: Acharya, Nirajan, et al.
Veröffentlicht: (2026)
FNF: Functional Network Fingerprint for Large Language Models
von: Liu, Yiheng, et al.
Veröffentlicht: (2026)
von: Liu, Yiheng, et al.
Veröffentlicht: (2026)
The Dark Side of Function Calling: Pathways to Jailbreaking Large Language Models
von: Wu, Zihui, et al.
Veröffentlicht: (2024)
von: Wu, Zihui, et al.
Veröffentlicht: (2024)
Certified Adversarial Robustness of Machine Learning-based Malware Detectors via (De)Randomized Smoothing
von: Gibert, Daniel, et al.
Veröffentlicht: (2024)
von: Gibert, Daniel, et al.
Veröffentlicht: (2024)
May I have your Attention? Breaking Fine-Tuning based Prompt Injection Defenses using Architecture-Aware Attacks
von: Pandya, Nishit V., et al.
Veröffentlicht: (2025)
von: Pandya, Nishit V., et al.
Veröffentlicht: (2025)
Securing Agentic AI: Threat Modeling and Risk Analysis for Network Monitoring Agentic AI System
von: Zambare, Pallavi, et al.
Veröffentlicht: (2025)
von: Zambare, Pallavi, et al.
Veröffentlicht: (2025)
Actionable Cyber Threat Intelligence using Knowledge Graphs and Large Language Models
von: Fieblinger, Romy, et al.
Veröffentlicht: (2024)
von: Fieblinger, Romy, et al.
Veröffentlicht: (2024)
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
von: Zhang, Dongsen, et al.
Veröffentlicht: (2025)
von: Zhang, Dongsen, et al.
Veröffentlicht: (2025)
The Use of Large Language Models (LLM) for Cyber Threat Intelligence (CTI) in Cybercrime Forums
von: Clairoux-Trepanier, Vanessa, et al.
Veröffentlicht: (2024)
von: Clairoux-Trepanier, Vanessa, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HarmLevelBench: Evaluating Harm-Level Compliance and the Impact of Quantization on Model Alignment
von: Belkhiter, Yannis, et al.
Veröffentlicht: (2024) -
Step-Tagging: Toward controlling the generation of Language Reasoning Models through step monitoring
von: Belkhiter, Yannis, et al.
Veröffentlicht: (2025) -
Blue Teaming Function-Calling Agents
von: Dolcetti, Greta, et al.
Veröffentlicht: (2026) -
TRACES: Tagging Reasoning Steps for Adaptive Cost-Efficient Early-Stopping
von: Belkhiter, Yannis, et al.
Veröffentlicht: (2026) -
Pre-Hoc Predictions in AutoML: Leveraging LLMs to Enhance Model Selection and Benchmarking for Tabular datasets
von: Belkhiter, Yannis, et al.
Veröffentlicht: (2025)