A Comparative Evaluation of AI Agent Security Guardrails

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Qi, Li, Jiu, Wei, Pingtao, Xu, Jianjun, Wei, Xueyi, Shi, Jiwei, Zhang, Xuan, Yang, Yanhui, Hui, Xiaodong, Xu, Peng, Zhou, Lingquan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913066753458176
author Li, Qi
Li, Jiu
Wei, Pingtao
Xu, Jianjun
Wei, Xueyi
Shi, Jiwei
Zhang, Xuan
Yang, Yanhui
Hui, Xiaodong
Xu, Peng
Zhou, Lingquan
author_facet Li, Qi
Li, Jiu
Wei, Pingtao
Xu, Jianjun
Wei, Xueyi
Shi, Jiwei
Zhang, Xuan
Yang, Yanhui
Hui, Xiaodong
Xu, Peng
Zhou, Lingquan
contents This report presents a comparative evaluation of DKnownAI Guard in AI agent security scenarios, benchmarked against three competing products: AWS Bedrock Guardrails, Azure Content Safety, and Lakera Guard. Using human annotation as the ground truth, we assess each guardrail's ability to detect two categories of risks: threats to the agent itself (e.g., instruction override, indirect injection, tool abuse) and requests intended to elicit harmful content (e.g., hate speech, pornography, violence). Evaluation results demonstrate that DKnownAI Guard achieves the highest recall rate at 96.5\% and ranks first in true negative rate (TNR) at 90.4\%, delivering the best overall performance among all evaluated guardrails.
format Preprint
id arxiv_https___arxiv_org_abs_2604_24826
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Comparative Evaluation of AI Agent Security Guardrails
Li, Qi
Li, Jiu
Wei, Pingtao
Xu, Jianjun
Wei, Xueyi
Shi, Jiwei
Zhang, Xuan
Yang, Yanhui
Hui, Xiaodong
Xu, Peng
Zhou, Lingquan
Cryptography and Security
Artificial Intelligence
This report presents a comparative evaluation of DKnownAI Guard in AI agent security scenarios, benchmarked against three competing products: AWS Bedrock Guardrails, Azure Content Safety, and Lakera Guard. Using human annotation as the ground truth, we assess each guardrail's ability to detect two categories of risks: threats to the agent itself (e.g., instruction override, indirect injection, tool abuse) and requests intended to elicit harmful content (e.g., hate speech, pornography, violence). Evaluation results demonstrate that DKnownAI Guard achieves the highest recall rate at 96.5\% and ranks first in true negative rate (TNR) at 90.4\%, delivering the best overall performance among all evaluated guardrails.
title A Comparative Evaluation of AI Agent Security Guardrails
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2604.24826