Agent Guide: A Simple Agent Behavioral Watermarking Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Kaibo, Zhang, Zipei, Yang, Zhongliang, Zhou, Linna |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robust Uncertainty Quantification for Factual Generation of Large Language Models
by: Zhang, Yuhao, et al.
Published: (2026)
by: Zhang, Yuhao, et al.
Published: (2026)
Cross-LLM Generalization of Behavioral Backdoor Detection in AI Agent Supply Chains
by: Sanna, Arun Chowdary
Published: (2025)
by: Sanna, Arun Chowdary
Published: (2025)
The Importance of Out-of-Band Metadata for Safe Autonomous Agents: The Redpanda Agentic Data Plane
by: Akidau, Tyler, et al.
Published: (2026)
by: Akidau, Tyler, et al.
Published: (2026)
Measuring Harmfulness of Computer-Using Agents
by: Tian, Aaron Xuxiang, et al.
Published: (2025)
by: Tian, Aaron Xuxiang, et al.
Published: (2025)
AgentMark: Utility-Preserving Behavioral Watermarking for Agents
by: Huang, Kaibo, et al.
Published: (2026)
by: Huang, Kaibo, et al.
Published: (2026)
Safe and Policy-Compliant Multi-Agent Orchestration for Enterprise AI
by: Pasupuleti, Vinil, et al.
Published: (2026)
by: Pasupuleti, Vinil, et al.
Published: (2026)
Security Attack and Defense Strategies for Autonomous Agent Frameworks: A Layered Review with OpenClaw as a Case Study
by: Xu, Luyao, et al.
Published: (2026)
by: Xu, Luyao, et al.
Published: (2026)
Benchmarking Autonomous Agents against Temporal, Spatial, and Semantic Evasions
by: Ma, Jianan, et al.
Published: (2026)
by: Ma, Jianan, et al.
Published: (2026)
AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purification
by: Zhang, Tian, et al.
Published: (2026)
by: Zhang, Tian, et al.
Published: (2026)
Spill The Beans: Exploiting CPU Cache Side-Channels to Leak Tokens from Large Language Models
by: Adiletta, Andrew, et al.
Published: (2025)
by: Adiletta, Andrew, et al.
Published: (2025)
Reverse CAPTCHA: Evaluating LLM Susceptibility to Invisible Unicode Instruction Injection
by: Graves, Marcus
Published: (2026)
by: Graves, Marcus
Published: (2026)
Detecting Sleeper Agents in Large Language Models via Semantic Drift Analysis
by: Zanbaghi, Shahin, et al.
Published: (2025)
by: Zanbaghi, Shahin, et al.
Published: (2025)
Right to History: A Sovereignty Kernel for Verifiable AI Agent Execution
by: Zhang, Jing
Published: (2026)
by: Zhang, Jing
Published: (2026)
HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing
by: Muzsai, Lajos, et al.
Published: (2024)
by: Muzsai, Lajos, et al.
Published: (2024)
The Open-Weight Paradox: Why Restricting Access to AI Models May Undermine the Safety It Seeks to Protect
by: Gomes, Vinicius Santana
Published: (2026)
by: Gomes, Vinicius Santana
Published: (2026)
HBEE: Human Behavioral Entropy Engine -- Pre-Registered Multi-Agent LLM Simulation of Peer-Suspicion-Based Detection Inversion
by: Ferrel, Vickson
Published: (2026)
by: Ferrel, Vickson
Published: (2026)
An Agentic Multi-Agent Architecture for Cybersecurity Risk Management
by: Gupta, Ravish, et al.
Published: (2026)
by: Gupta, Ravish, et al.
Published: (2026)
Governance Architecture for Autonomous Agent Systems: Threats, Framework, and Engineering Practice
by: Ge, Yuxu
Published: (2026)
by: Ge, Yuxu
Published: (2026)
Criminal Liability in AI-Enabled Autonomous Vehicles: A Comparative Study
by: Singh, Sahibpreet, et al.
Published: (2025)
by: Singh, Sahibpreet, et al.
Published: (2025)
Towards Low-Latency and Adaptive Ransomware Detection Using Contrastive Learning
by: Pan, Zhixin, et al.
Published: (2025)
by: Pan, Zhixin, et al.
Published: (2025)
Architecture-Agnostic Feature Synergy for Universal Defense Against Heterogeneous Generative Threats
by: Zhang, Bingxue, et al.
Published: (2026)
by: Zhang, Bingxue, et al.
Published: (2026)
Cybercrime and Computer Forensics in Epoch of Artificial Intelligence in India
by: Singh, Sahibpreet, et al.
Published: (2025)
by: Singh, Sahibpreet, et al.
Published: (2025)
A Survey on the Security of Long-Term Memory in LLM Agents: Toward Mnemonic Sovereignty
by: Lin, Zehao, et al.
Published: (2026)
by: Lin, Zehao, et al.
Published: (2026)
Improving LLM Agents with Reinforcement Learning on Cryptographic CTF Challenges
by: Muzsai, Lajos, et al.
Published: (2025)
by: Muzsai, Lajos, et al.
Published: (2025)
From nuclear safety to LLM security: Applying non-probabilistic risk management strategies to build safe and secure LLM-powered systems
by: Gutfraind, Alexander, et al.
Published: (2025)
by: Gutfraind, Alexander, et al.
Published: (2025)
Do Latent Tokens Think? A Causal and Adversarial Analysis of Chain-of-Continuous-Thought
by: Zhang, Yuyi, et al.
Published: (2025)
by: Zhang, Yuyi, et al.
Published: (2025)
Unveiling the Vulnerability of Private Fine-Tuning in Split-Based Frameworks for Large Language Models: A Bidirectionally Enhanced Attack
by: Chen, Guanzhong, et al.
Published: (2024)
by: Chen, Guanzhong, et al.
Published: (2024)
Hidden Reliability Risks in Large Language Models: Systematic Identification of Precision-Induced Output Disagreements
by: Wang, Yifei, et al.
Published: (2026)
by: Wang, Yifei, et al.
Published: (2026)
Refute-or-Promote: An Adversarial Stage-Gated Multi-Agent Review Methodology for High-Precision LLM-Assisted Defect Discovery
by: Agarwal, Abhinav
Published: (2026)
by: Agarwal, Abhinav
Published: (2026)
Countermind: A Multi-Layered Security Architecture for Large Language Models
by: Schwarz, Dominik
Published: (2025)
by: Schwarz, Dominik
Published: (2025)
Peak + Accumulation: A Proxy-Level Scoring Formula for Multi-Turn LLM Attack Detection
by: Corll, J Alex
Published: (2026)
by: Corll, J Alex
Published: (2026)
Depth-Dependent Indirect Prompt Injection in Tool-Calling ReAct Agents: Injection Depth, Payload Framing, and Turn-Budget Sensitivity
by: Rashidi, Mohammadreza
Published: (2026)
by: Rashidi, Mohammadreza
Published: (2026)
DeepSignature: Digitally Signed, Content-Encoding Watermarks for Robust and Transparent Image Authentication
by: Graf, Mathias, et al.
Published: (2026)
by: Graf, Mathias, et al.
Published: (2026)
A Dual-Path Generative Framework for Zero-Day Fraud Detection in Banking Systems
by: Ismail, Nasim Abdirahman, et al.
Published: (2026)
by: Ismail, Nasim Abdirahman, et al.
Published: (2026)
Send to which account? Evaluation of an LLM-based Scambaiting System
by: Siadati, Hossein, et al.
Published: (2025)
by: Siadati, Hossein, et al.
Published: (2025)
Design Principles for the Construction of a Benchmark Evaluating Security Operation Capabilities of Multi-agent AI Systems
by: Cai, Yicheng, et al.
Published: (2026)
by: Cai, Yicheng, et al.
Published: (2026)
Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps
by: Chona, Alankrit, et al.
Published: (2026)
by: Chona, Alankrit, et al.
Published: (2026)
Context Kubernetes: Declarative Orchestration of Enterprise Knowledge for Agentic AI Systems
by: Mouzouni, Charafeddine
Published: (2026)
by: Mouzouni, Charafeddine
Published: (2026)
Rescuing the Unpoisoned: Efficient Defense against Knowledge Corruption Attacks on RAG Systems
by: Kim, Minseok, et al.
Published: (2025)
by: Kim, Minseok, et al.
Published: (2025)
Closing the Distribution Gap in Adversarial Training for LLMs
by: Hu, Chengzhi, et al.
Published: (2026)
by: Hu, Chengzhi, et al.
Published: (2026)
Similar Items
-
Robust Uncertainty Quantification for Factual Generation of Large Language Models
by: Zhang, Yuhao, et al.
Published: (2026) -
Cross-LLM Generalization of Behavioral Backdoor Detection in AI Agent Supply Chains
by: Sanna, Arun Chowdary
Published: (2025) -
The Importance of Out-of-Band Metadata for Safe Autonomous Agents: The Redpanda Agentic Data Plane
by: Akidau, Tyler, et al.
Published: (2026) -
Measuring Harmfulness of Computer-Using Agents
by: Tian, Aaron Xuxiang, et al.
Published: (2025) -
AgentMark: Utility-Preserving Behavioral Watermarking for Agents
by: Huang, Kaibo, et al.
Published: (2026)