QuadSentinel: Sequent Safety for Machine-Checkable Control in Multi-agent Systems
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912773263327232 |
|---|---|
| author | Yang, Yiliu Jiang, Yilei Wang, Qunzhong Tan, Yingshui Zhu, Xiaoyong Chow, Sherman S. M. Zheng, Bo Yue, Xiangyu |
| author_facet | Yang, Yiliu Jiang, Yilei Wang, Qunzhong Tan, Yingshui Zhu, Xiaoyong Chow, Sherman S. M. Zheng, Bo Yue, Xiangyu |
| contents | Safety risks arise as large language model-based agents solve complex tasks with tools, multi-step plans, and inter-agent messages. However, deployer-written policies in natural language are ambiguous and context dependent, so they map poorly to machine-checkable rules, and runtime enforcement is unreliable. Expressing safety policies as sequents, we propose \textsc{QuadSentinel}, a four-agent guard (state tracker, policy verifier, threat watcher, and referee) that compiles these policies into machine-checkable rules built from predicates over observable state and enforces them online. Referee logic plus an efficient top-$k$ predicate updater keeps costs low by prioritizing checks and resolving conflicts hierarchically. Measured on ST-WebAgentBench (ICML CUA~'25) and AgentHarm (ICLR~'25), \textsc{QuadSentinel} improves guardrail accuracy and rule recall while reducing false positives. Against single-agent baselines such as ShieldAgent (ICML~'25), it yields better overall safety control. Near-term deployments can adopt this pattern without modifying core agents by keeping policies separate and machine-checkable. Our code will be made publicly available at https://github.com/yyiliu/QuadSentinel. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_16279 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | QuadSentinel: Sequent Safety for Machine-Checkable Control in Multi-agent Systems Yang, Yiliu Jiang, Yilei Wang, Qunzhong Tan, Yingshui Zhu, Xiaoyong Chow, Sherman S. M. Zheng, Bo Yue, Xiangyu Artificial Intelligence Computation and Language Safety risks arise as large language model-based agents solve complex tasks with tools, multi-step plans, and inter-agent messages. However, deployer-written policies in natural language are ambiguous and context dependent, so they map poorly to machine-checkable rules, and runtime enforcement is unreliable. Expressing safety policies as sequents, we propose \textsc{QuadSentinel}, a four-agent guard (state tracker, policy verifier, threat watcher, and referee) that compiles these policies into machine-checkable rules built from predicates over observable state and enforces them online. Referee logic plus an efficient top-$k$ predicate updater keeps costs low by prioritizing checks and resolving conflicts hierarchically. Measured on ST-WebAgentBench (ICML CUA~'25) and AgentHarm (ICLR~'25), \textsc{QuadSentinel} improves guardrail accuracy and rule recall while reducing false positives. Against single-agent baselines such as ShieldAgent (ICML~'25), it yields better overall safety control. Near-term deployments can adopt this pattern without modifying core agents by keeping policies separate and machine-checkable. Our code will be made publicly available at https://github.com/yyiliu/QuadSentinel. |
| title | QuadSentinel: Sequent Safety for Machine-Checkable Control in Multi-agent Systems |
| topic | Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2512.16279 |