Learning Efficient Guardrails for Compliance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wen, Xiaofei, Mo, Wenjie Jacky, Xie, Yanan, Qi, Peng, Chen, Muhao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917508618911744
author Wen, Xiaofei
Mo, Wenjie Jacky
Xie, Yanan
Qi, Peng
Chen, Muhao
author_facet Wen, Xiaofei
Mo, Wenjie Jacky
Xie, Yanan
Qi, Peng
Chen, Muhao
contents Autonomous web agents are increasingly deployed for long-horizon tasks, yet their ability to adhere to real-world policies remains critically underexplored compared to standard safety objectives. To address this gap, we introduce PolicyGuardBench, a benchmark of 60k policy-trajectory pairs designed to evaluate compliance through both full-trajectory and novel prefix-based violation detection tasks. Using this dataset, we train PolicyGuard, a lightweight guardrail model that achieves strong detection accuracy while maintaining high inference efficiency. Notably, our model demonstrates robust generalization capabilities, preserving high performance even on unseen domains. These contributions establish a comprehensive framework for studying policy compliance, showing that accurate and generalizable guardrails are feasible at small scales.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03485
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Efficient Guardrails for Compliance
Wen, Xiaofei
Mo, Wenjie Jacky
Xie, Yanan
Qi, Peng
Chen, Muhao
Artificial Intelligence
I.2.7
Autonomous web agents are increasingly deployed for long-horizon tasks, yet their ability to adhere to real-world policies remains critically underexplored compared to standard safety objectives. To address this gap, we introduce PolicyGuardBench, a benchmark of 60k policy-trajectory pairs designed to evaluate compliance through both full-trajectory and novel prefix-based violation detection tasks. Using this dataset, we train PolicyGuard, a lightweight guardrail model that achieves strong detection accuracy while maintaining high inference efficiency. Notably, our model demonstrates robust generalization capabilities, preserving high performance even on unseen domains. These contributions establish a comprehensive framework for studying policy compliance, showing that accurate and generalizable guardrails are feasible at small scales.
title Learning Efficient Guardrails for Compliance
topic Artificial Intelligence
I.2.7
url https://arxiv.org/abs/2510.03485