We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Fang, Junfeng, Yao, Zijun, Wang, Ruipeng, Ma, Haokai, Wang, Xiang, Chua, Tat-Seng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Models
by: Fang, Junfeng, et al.
Published: (2025)
by: Fang, Junfeng, et al.
Published: (2025)
TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety
by: Hong, Zhepei, et al.
Published: (2026)
by: Hong, Zhepei, et al.
Published: (2026)
AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition
by: Wang, Ruipeng, et al.
Published: (2026)
by: Wang, Ruipeng, et al.
Published: (2026)
DRAFT: Task Decoupled Latent Reasoning for Agent Safety
by: Wang, Lin, et al.
Published: (2026)
by: Wang, Lin, et al.
Published: (2026)
DualEdit: Mitigating Safety Fallback in LLM Backdoor Editing via Affirmation-Refusal Regulation
by: Jiang, Houcheng, et al.
Published: (2025)
by: Jiang, Houcheng, et al.
Published: (2025)
On Generative Agents in Recommendation
by: Zhang, An, et al.
Published: (2023)
by: Zhang, An, et al.
Published: (2023)
Disentangling Masked Autoencoders for Unsupervised Domain Generalization
by: Zhang, An, et al.
Published: (2024)
by: Zhang, An, et al.
Published: (2024)
Hello Again! LLM-powered Personalized Agent for Long-term Dialogue
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
LLM2Rec: Large Language Models Are Powerful Embedding Models for Sequential Recommendation
by: He, Yingzhi, et al.
Published: (2025)
by: He, Yingzhi, et al.
Published: (2025)
Plug-and-Play Policy Planner for Large Language Model Powered Dialogue Agents
by: Deng, Yang, et al.
Published: (2023)
by: Deng, Yang, et al.
Published: (2023)
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
by: Jin, Zhe, et al.
Published: (2025)
by: Jin, Zhe, et al.
Published: (2025)
NExT-Guard: Training-Free Streaming Safeguard without Token-Level Labels
by: Fang, Junfeng, et al.
Published: (2026)
by: Fang, Junfeng, et al.
Published: (2026)
Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments
by: Chen, Yuxin, et al.
Published: (2026)
by: Chen, Yuxin, et al.
Published: (2026)
Risky-Bench: Probing Agentic Safety Risks under Real-World Deployment
by: Zheng, Jingnan, et al.
Published: (2026)
by: Zheng, Jingnan, et al.
Published: (2026)
Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics
by: Sheng, Leheng, et al.
Published: (2026)
by: Sheng, Leheng, et al.
Published: (2026)
Towards Goal-oriented Intelligent Tutoring Systems in Online Education
by: Deng, Yang, et al.
Published: (2023)
by: Deng, Yang, et al.
Published: (2023)
AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint
by: Sheng, Leheng, et al.
Published: (2025)
by: Sheng, Leheng, et al.
Published: (2025)
ALI-Agent: Assessing LLMs' Alignment with Human Values via Agent-based Evaluation
by: Zheng, Jingnan, et al.
Published: (2024)
by: Zheng, Jingnan, et al.
Published: (2024)
AlphaEdit: Null-Space Constrained Knowledge Editing for Language Models
by: Fang, Junfeng, et al.
Published: (2024)
by: Fang, Junfeng, et al.
Published: (2024)
Can I Trust Your Answer? Visually Grounded Video Question Answering
by: Xiao, Junbin, et al.
Published: (2023)
by: Xiao, Junbin, et al.
Published: (2023)
MCP-ITP: An Automated Framework for Implicit Tool Poisoning in MCP
by: Li, Ruiqi, et al.
Published: (2026)
by: Li, Ruiqi, et al.
Published: (2026)
Language Representations Can be What Recommenders Need: Findings and Potentials
by: Sheng, Leheng, et al.
Published: (2024)
by: Sheng, Leheng, et al.
Published: (2024)
Enhancing Multi-Agent Consensus through Third-Party LLM Integration: Analyzing Uncertainty and Mitigating Hallucinations in Large Language Models
by: Duan, Zhihua, et al.
Published: (2024)
by: Duan, Zhihua, et al.
Published: (2024)
AgentTrap: Measuring Runtime Trust Failures in Third-Party Agent Skills
by: Zhuang, Haomin, et al.
Published: (2026)
by: Zhuang, Haomin, et al.
Published: (2026)
Rubric-based On-policy Distillation
by: Fang, Junfeng, et al.
Published: (2026)
by: Fang, Junfeng, et al.
Published: (2026)
Active Zero: Self-Evolving Vision-Language Models through Active Environment Exploration
by: He, Jinghan, et al.
Published: (2026)
by: He, Jinghan, et al.
Published: (2026)
Ask-before-Plan: Proactive Language Agents for Real-World Planning
by: Zhang, Xuan, et al.
Published: (2024)
by: Zhang, Xuan, et al.
Published: (2024)
Mitigating the Negative Impact of Over-association for Conversational Query Production
by: Wang, Ante, et al.
Published: (2024)
by: Wang, Ante, et al.
Published: (2024)
Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
by: Li, Gengsheng, et al.
Published: (2026)
by: Li, Gengsheng, et al.
Published: (2026)
Towards Temporal-Aware Multi-Modal Retrieval Augmented Generation in Finance
by: Zhu, Fengbin, et al.
Published: (2025)
by: Zhu, Fengbin, et al.
Published: (2025)
Beyond Persuasion: Towards Conversational Recommender System with Credible Explanations
by: Qin, Peixin, et al.
Published: (2024)
by: Qin, Peixin, et al.
Published: (2024)
FashionReGen: LLM-Empowered Fashion Report Generation
by: Ding, Yujuan, et al.
Published: (2024)
by: Ding, Yujuan, et al.
Published: (2024)
Large Language Models Empowered Personalized Web Agents
by: Cai, Hongru, et al.
Published: (2024)
by: Cai, Hongru, et al.
Published: (2024)
SafeMind: Benchmarking and Mitigating Safety Risks in Embodied LLM Agents
by: Chen, Ruolin, et al.
Published: (2025)
by: Chen, Ruolin, et al.
Published: (2025)
NextMem: Towards Latent Factual Memory for LLM-based Agents
by: Zhang, Zeyu, et al.
Published: (2026)
by: Zhang, Zeyu, et al.
Published: (2026)
A Survey on Neural Question Generation: Methods, Applications, and Prospects
by: Guo, Shasha, et al.
Published: (2024)
by: Guo, Shasha, et al.
Published: (2024)
MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems
by: Chen, Kai, et al.
Published: (2025)
by: Chen, Kai, et al.
Published: (2025)
SafeNeuron: Neuron-Level Safety Alignment for Large Language Models
by: Wang, Zhaoxin, et al.
Published: (2026)
by: Wang, Zhaoxin, et al.
Published: (2026)
On the Multi-turn Instruction Following for Conversational Web Agents
by: Deng, Yang, et al.
Published: (2024)
by: Deng, Yang, et al.
Published: (2024)
VisualTrap: A Stealthy Backdoor Attack on GUI Agents via Visual Grounding Manipulation
by: Ye, Ziang, et al.
Published: (2025)
by: Ye, Ziang, et al.
Published: (2025)
Similar Items
-
SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Models
by: Fang, Junfeng, et al.
Published: (2025) -
TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety
by: Hong, Zhepei, et al.
Published: (2026) -
AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition
by: Wang, Ruipeng, et al.
Published: (2026) -
DRAFT: Task Decoupled Latent Reasoning for Agent Safety
by: Wang, Lin, et al.
Published: (2026) -
DualEdit: Mitigating Safety Fallback in LLM Backdoor Editing via Affirmation-Refusal Regulation
by: Jiang, Houcheng, et al.
Published: (2025)