DynaGuard: A Dynamic Guardian Model With User-Defined Policies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hoover, Monte, Baherwani, Vatsal, Jain, Neel, Saifullah, Khalid, Vincent, Joseph, Jain, Chirag, Rad, Melissa Kazemi, Bruss, C. Bayan, Panda, Ashwinee, Goldstein, Tom
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912632312692736
author Hoover, Monte
Baherwani, Vatsal
Jain, Neel
Saifullah, Khalid
Vincent, Joseph
Jain, Chirag
Rad, Melissa Kazemi
Bruss, C. Bayan
Panda, Ashwinee
Goldstein, Tom
author_facet Hoover, Monte
Baherwani, Vatsal
Jain, Neel
Saifullah, Khalid
Vincent, Joseph
Jain, Chirag
Rad, Melissa Kazemi
Bruss, C. Bayan
Panda, Ashwinee
Goldstein, Tom
contents Guardian models play a crucial role in ensuring the safety and ethical behavior of user-facing AI applications by enforcing guardrails and detecting harmful content. While standard guardian models are limited to predefined, static harm categories, we introduce DynaGuard, a suite of dynamic guardian models offering novel flexibility by evaluating text based on user-defined policies, and DynaBench, a dataset for training and evaluating dynamic guardian models. Our models provide both rapid detection of policy violations and a chain-of-thought reasoning option that articulate and justify model outputs. Critically, DynaGuard not only surpasses static models in detection accuracy on traditional safety categories, but is competitive with frontier reasoning models on free-form policy violations, all in a fraction of the time. This makes DynaGuard an critical tool for language model guardrails.
format Preprint
id arxiv_https___arxiv_org_abs_2509_02563
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DynaGuard: A Dynamic Guardian Model With User-Defined Policies
Hoover, Monte
Baherwani, Vatsal
Jain, Neel
Saifullah, Khalid
Vincent, Joseph
Jain, Chirag
Rad, Melissa Kazemi
Bruss, C. Bayan
Panda, Ashwinee
Goldstein, Tom
Machine Learning
Computation and Language
Guardian models play a crucial role in ensuring the safety and ethical behavior of user-facing AI applications by enforcing guardrails and detecting harmful content. While standard guardian models are limited to predefined, static harm categories, we introduce DynaGuard, a suite of dynamic guardian models offering novel flexibility by evaluating text based on user-defined policies, and DynaBench, a dataset for training and evaluating dynamic guardian models. Our models provide both rapid detection of policy violations and a chain-of-thought reasoning option that articulate and justify model outputs. Critically, DynaGuard not only surpasses static models in detection accuracy on traditional safety categories, but is competitive with frontier reasoning models on free-form policy violations, all in a fraction of the time. This makes DynaGuard an critical tool for language model guardrails.
title DynaGuard: A Dynamic Guardian Model With User-Defined Policies
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2509.02563