Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qi, Jinhu, Li, Muzhi, Liu, Jiahong, Shu, Yuqin, Yu, Dianzhi, Ma, Shicheng, Cui, Wenqian, Zhao, Yiyang, Chen, Yiyi, Jiang, Ruoxi, King, Irwin, Xu, Zenglin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917526172073984
author Qi, Jinhu
Li, Muzhi
Liu, Jiahong
Shu, Yuqin
Yu, Dianzhi
Ma, Shicheng
Cui, Wenqian
Zhao, Yiyang
Chen, Yiyi
Jiang, Ruoxi
King, Irwin
Xu, Zenglin
author_facet Qi, Jinhu
Li, Muzhi
Liu, Jiahong
Shu, Yuqin
Yu, Dianzhi
Ma, Shicheng
Cui, Wenqian
Zhao, Yiyang
Chen, Yiyi
Jiang, Ruoxi
King, Irwin
Xu, Zenglin
contents Agentic AI systems -- Large Language Models (LLMs) augmented with planning, tool use, memory, and long-horizon interactions -- can execute complex tasks autonomously, but their multi-step trajectories introduce new failure modes that challenge trustworthiness. This survey provides a focused examination of trustworthy agentic AI through two core dimensions that are critical for high-risk deployments: Safety and Robustness, and Privacy and System Security. For each dimension, we clarify key concepts, identify where risks emerge along the agent workflow, and summarize stage-targeted mitigation strategies. Other trustworthiness aspects (value alignment, transparency, fairness, and accountability) are discussed as relevant context rather than parallel chapters. To support consistent comparison and deployment decisions, we consolidate evaluation into a unified metrics-and-benchmarks hub, emphasizing both outcome and process signals (e.g., constraint violations, trace completeness, and adversarial success rates) and offering scenario-to-metric guidance for release gating. We conclude by outlining open challenges such as self-evolving agents, runtime monitoring and verification, privacy-preserving personalization, and the trust-utility trade-off, and present a case study of real-world security failures in open-source agentic systems. Our goal is to serve as a practical reference for researchers and practitioners building trustworthy agentic systems in high-stakes environments.
format Preprint
id arxiv_https___arxiv_org_abs_2605_23989
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security
Qi, Jinhu
Li, Muzhi
Liu, Jiahong
Shu, Yuqin
Yu, Dianzhi
Ma, Shicheng
Cui, Wenqian
Zhao, Yiyang
Chen, Yiyi
Jiang, Ruoxi
King, Irwin
Xu, Zenglin
Artificial Intelligence
Computation and Language
Cryptography and Security
I.2.11; K.6.5
Agentic AI systems -- Large Language Models (LLMs) augmented with planning, tool use, memory, and long-horizon interactions -- can execute complex tasks autonomously, but their multi-step trajectories introduce new failure modes that challenge trustworthiness. This survey provides a focused examination of trustworthy agentic AI through two core dimensions that are critical for high-risk deployments: Safety and Robustness, and Privacy and System Security. For each dimension, we clarify key concepts, identify where risks emerge along the agent workflow, and summarize stage-targeted mitigation strategies. Other trustworthiness aspects (value alignment, transparency, fairness, and accountability) are discussed as relevant context rather than parallel chapters. To support consistent comparison and deployment decisions, we consolidate evaluation into a unified metrics-and-benchmarks hub, emphasizing both outcome and process signals (e.g., constraint violations, trace completeness, and adversarial success rates) and offering scenario-to-metric guidance for release gating. We conclude by outlining open challenges such as self-evolving agents, runtime monitoring and verification, privacy-preserving personalization, and the trust-utility trade-off, and present a case study of real-world security failures in open-source agentic systems. Our goal is to serve as a practical reference for researchers and practitioners building trustworthy agentic systems in high-stakes environments.
title Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security
topic Artificial Intelligence
Computation and Language
Cryptography and Security
I.2.11; K.6.5
url https://arxiv.org/abs/2605.23989