Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Imperial, Joseph Marvin, Madabushi, Harish Tayyar
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914061104447488
author Imperial, Joseph Marvin
Madabushi, Harish Tayyar
author_facet Imperial, Joseph Marvin
Madabushi, Harish Tayyar
contents Policy compliance assessment is a fundamental task of evaluating whether an input case strictly complies with a set of human-defined rules, more generally known as policies. In practice, human experts follow a systematic, step-by-step process to identify violations with respect to specific stipulations outlined in the policy. However, such documentation of gold-standard, expert-level reasoning processes is costly to acquire. In this paper, we introduce Policy Reasoning Traces (PRT), a form of specialized generated reasoning chains that serve as a reasoning bridge to improve an LLM's policy compliance assessment capabilities. Our empirical evaluations demonstrate that the use of PRTs for both inference-time and training-time scenarios significantly enhances the performance of open-weight and commercial models, setting a new state-of-the-art for HIPAA and GDPR policies. Beyond accuracy gains, we also highlight how PRTs can improve an LLM's ability to accurately cite policy clauses, as well as influence compliance decisions through their high utilization from the raw chains of thought.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23291
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces
Imperial, Joseph Marvin
Madabushi, Harish Tayyar
Computation and Language
Machine Learning
Policy compliance assessment is a fundamental task of evaluating whether an input case strictly complies with a set of human-defined rules, more generally known as policies. In practice, human experts follow a systematic, step-by-step process to identify violations with respect to specific stipulations outlined in the policy. However, such documentation of gold-standard, expert-level reasoning processes is costly to acquire. In this paper, we introduce Policy Reasoning Traces (PRT), a form of specialized generated reasoning chains that serve as a reasoning bridge to improve an LLM's policy compliance assessment capabilities. Our empirical evaluations demonstrate that the use of PRTs for both inference-time and training-time scenarios significantly enhances the performance of open-weight and commercial models, setting a new state-of-the-art for HIPAA and GDPR policies. Beyond accuracy gains, we also highlight how PRTs can improve an LLM's ability to accurately cite policy clauses, as well as influence compliance decisions through their high utilization from the raw chains of thought.
title Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2509.23291