Agent Guide: A Simple Agent Behavioral Watermarking Framework

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Kaibo, Zhang, Zipei, Yang, Zhongliang, Zhou, Linna
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918534078005248
author Huang, Kaibo
Zhang, Zipei
Yang, Zhongliang
Zhou, Linna
author_facet Huang, Kaibo
Zhang, Zipei
Yang, Zhongliang
Zhou, Linna
contents The increasing deployment of intelligent agents in digital ecosystems, such as social media platforms, has raised significant concerns about traceability and accountability, particularly in cybersecurity and digital content protection. Traditional large language model (LLM) watermarking techniques, which rely on token-level manipulations, are ill-suited for agents due to the challenges of behavior tokenization and information loss during behavior-to-action translation. To address these issues, we propose Agent Guide, a novel behavioral watermarking framework that embeds watermarks by guiding the agent's high-level decisions (behavior) through probability biases, while preserving the naturalness of specific executions (action). Our approach decouples agent behavior into two levels, behavior (e.g., choosing to bookmark) and action (e.g., bookmarking with specific tags), and applies watermark-guided biases to the behavior probability distribution. We employ a z-statistic-based statistical analysis to detect the watermark, ensuring reliable extraction over multiple rounds. Experiments in a social media scenario with diverse agent profiles demonstrate that Agent Guide achieves effective watermark detection with a low false positive rate. Our framework provides a practical and robust solution for agent watermarking, with applications in identifying malicious agents and protecting proprietary agent systems.
format Preprint
id arxiv_https___arxiv_org_abs_2504_05871
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Agent Guide: A Simple Agent Behavioral Watermarking Framework
Huang, Kaibo
Zhang, Zipei
Yang, Zhongliang
Zhou, Linna
Artificial Intelligence
K.6.5
The increasing deployment of intelligent agents in digital ecosystems, such as social media platforms, has raised significant concerns about traceability and accountability, particularly in cybersecurity and digital content protection. Traditional large language model (LLM) watermarking techniques, which rely on token-level manipulations, are ill-suited for agents due to the challenges of behavior tokenization and information loss during behavior-to-action translation. To address these issues, we propose Agent Guide, a novel behavioral watermarking framework that embeds watermarks by guiding the agent's high-level decisions (behavior) through probability biases, while preserving the naturalness of specific executions (action). Our approach decouples agent behavior into two levels, behavior (e.g., choosing to bookmark) and action (e.g., bookmarking with specific tags), and applies watermark-guided biases to the behavior probability distribution. We employ a z-statistic-based statistical analysis to detect the watermark, ensuring reliable extraction over multiple rounds. Experiments in a social media scenario with diverse agent profiles demonstrate that Agent Guide achieves effective watermark detection with a low false positive rate. Our framework provides a practical and robust solution for agent watermarking, with applications in identifying malicious agents and protecting proprietary agent systems.
title Agent Guide: A Simple Agent Behavioral Watermarking Framework
topic Artificial Intelligence
K.6.5
url https://arxiv.org/abs/2504.05871