An Evaluation Framework for LLM-Driven Regulatory-to-Policy-as-Code Translation - Replication Package

Fuente: Zenodo
Salvato in:
Dettagli Bibliografici
Autore principale: Nguyen, Bao Van
Natura: Recurso digital
Pubblicazione: Zenodo 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866902121026158592
author Nguyen, Bao Van
author_facet Nguyen, Bao Van
contents <p>Replication package for the evaluation framework. Contains: (1) a four-layer LLM-driven regulatory-to-policy-as-code pipeline (input → RAG retrieval → two-stage LLM generation → automated validation); (2) the Policy Accuracy Score (PAS) rubric implementation; (3) a 30-scenario NIST SP 800-53–derived dataset with reference OPA/Rego policies, reference Terraform configurations, and paired compliant / non-compliant OPA test fixtures; (4) the 180-run experiment that produced the results reported in the paper (3 commercial LLMs × 2 generation methods × 30 scenarios); (5) cached model outputs and OPA / Checkov proof artefacts enabling offline re-scoring without API access; (6) the paper-ready figures and tables generated from the canonical results CSV. The canonical results CSV has SHA-256 42f6ec840d7745653749d7a39aea6e338afbd81e30b8f5df7099b31e1de2643c.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_20399040
institution Zenodo
language
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle An Evaluation Framework for LLM-Driven Regulatory-to-Policy-as-Code Translation - Replication Package
Nguyen, Bao Van
large language models
evaluation framework
DevSecOps
policy as code
regulatory compliance
retrieval-augmented generation
NIST SP 800-53
Open Policy Agent
Rego
Terraform
<p>Replication package for the evaluation framework. Contains: (1) a four-layer LLM-driven regulatory-to-policy-as-code pipeline (input → RAG retrieval → two-stage LLM generation → automated validation); (2) the Policy Accuracy Score (PAS) rubric implementation; (3) a 30-scenario NIST SP 800-53–derived dataset with reference OPA/Rego policies, reference Terraform configurations, and paired compliant / non-compliant OPA test fixtures; (4) the 180-run experiment that produced the results reported in the paper (3 commercial LLMs × 2 generation methods × 30 scenarios); (5) cached model outputs and OPA / Checkov proof artefacts enabling offline re-scoring without API access; (6) the paper-ready figures and tables generated from the canonical results CSV. The canonical results CSV has SHA-256 42f6ec840d7745653749d7a39aea6e338afbd81e30b8f5df7099b31e1de2643c.</p>
title An Evaluation Framework for LLM-Driven Regulatory-to-Policy-as-Code Translation - Replication Package
topic large language models
evaluation framework
DevSecOps
policy as code
regulatory compliance
retrieval-augmented generation
NIST SP 800-53
Open Policy Agent
Rego
Terraform
url https://doi.org/10.5281/zenodo.20399040