QuadSentinel: Sequent Safety for Machine-Checkable Control in Multi-agent Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Yiliu, Jiang, Yilei, Wang, Qunzhong, Tan, Yingshui, Zhu, Xiaoyong, Chow, Sherman S. M., Zheng, Bo, Yue, Xiangyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912773263327232
author Yang, Yiliu
Jiang, Yilei
Wang, Qunzhong
Tan, Yingshui
Zhu, Xiaoyong
Chow, Sherman S. M.
Zheng, Bo
Yue, Xiangyu
author_facet Yang, Yiliu
Jiang, Yilei
Wang, Qunzhong
Tan, Yingshui
Zhu, Xiaoyong
Chow, Sherman S. M.
Zheng, Bo
Yue, Xiangyu
contents Safety risks arise as large language model-based agents solve complex tasks with tools, multi-step plans, and inter-agent messages. However, deployer-written policies in natural language are ambiguous and context dependent, so they map poorly to machine-checkable rules, and runtime enforcement is unreliable. Expressing safety policies as sequents, we propose \textsc{QuadSentinel}, a four-agent guard (state tracker, policy verifier, threat watcher, and referee) that compiles these policies into machine-checkable rules built from predicates over observable state and enforces them online. Referee logic plus an efficient top-$k$ predicate updater keeps costs low by prioritizing checks and resolving conflicts hierarchically. Measured on ST-WebAgentBench (ICML CUA~'25) and AgentHarm (ICLR~'25), \textsc{QuadSentinel} improves guardrail accuracy and rule recall while reducing false positives. Against single-agent baselines such as ShieldAgent (ICML~'25), it yields better overall safety control. Near-term deployments can adopt this pattern without modifying core agents by keeping policies separate and machine-checkable. Our code will be made publicly available at https://github.com/yyiliu/QuadSentinel.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16279
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle QuadSentinel: Sequent Safety for Machine-Checkable Control in Multi-agent Systems
Yang, Yiliu
Jiang, Yilei
Wang, Qunzhong
Tan, Yingshui
Zhu, Xiaoyong
Chow, Sherman S. M.
Zheng, Bo
Yue, Xiangyu
Artificial Intelligence
Computation and Language
Safety risks arise as large language model-based agents solve complex tasks with tools, multi-step plans, and inter-agent messages. However, deployer-written policies in natural language are ambiguous and context dependent, so they map poorly to machine-checkable rules, and runtime enforcement is unreliable. Expressing safety policies as sequents, we propose \textsc{QuadSentinel}, a four-agent guard (state tracker, policy verifier, threat watcher, and referee) that compiles these policies into machine-checkable rules built from predicates over observable state and enforces them online. Referee logic plus an efficient top-$k$ predicate updater keeps costs low by prioritizing checks and resolving conflicts hierarchically. Measured on ST-WebAgentBench (ICML CUA~'25) and AgentHarm (ICLR~'25), \textsc{QuadSentinel} improves guardrail accuracy and rule recall while reducing false positives. Against single-agent baselines such as ShieldAgent (ICML~'25), it yields better overall safety control. Near-term deployments can adopt this pattern without modifying core agents by keeping policies separate and machine-checkable. Our code will be made publicly available at https://github.com/yyiliu/QuadSentinel.
title QuadSentinel: Sequent Safety for Machine-Checkable Control in Multi-agent Systems
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2512.16279