An Evaluation Framework for LLM-Driven Regulatory-to-Policy-as-Code Translation - Replication Package
Fuente:
Zenodo
Salvato in:
| Autore principale: | |
|---|---|
| Natura: | Recurso digital |
| Pubblicazione: |
Zenodo
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866902121026158592 |
|---|---|
| author | Nguyen, Bao Van |
| author_facet | Nguyen, Bao Van |
| contents | <p>Replication package for the evaluation framework. Contains: (1) a four-layer LLM-driven regulatory-to-policy-as-code pipeline (input → RAG retrieval → two-stage LLM generation → automated validation); (2) the Policy Accuracy Score (PAS) rubric implementation; (3) a 30-scenario NIST SP 800-53–derived dataset with reference OPA/Rego policies, reference Terraform configurations, and paired compliant / non-compliant OPA test fixtures; (4) the 180-run experiment that produced the results reported in the paper (3 commercial LLMs × 2 generation methods × 30 scenarios); (5) cached model outputs and OPA / Checkov proof artefacts enabling offline re-scoring without API access; (6) the paper-ready figures and tables generated from the canonical results CSV. The canonical results CSV has SHA-256 42f6ec840d7745653749d7a39aea6e338afbd81e30b8f5df7099b31e1de2643c.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_20399040 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | An Evaluation Framework for LLM-Driven Regulatory-to-Policy-as-Code Translation - Replication Package Nguyen, Bao Van large language models evaluation framework DevSecOps policy as code regulatory compliance retrieval-augmented generation NIST SP 800-53 Open Policy Agent Rego Terraform <p>Replication package for the evaluation framework. Contains: (1) a four-layer LLM-driven regulatory-to-policy-as-code pipeline (input → RAG retrieval → two-stage LLM generation → automated validation); (2) the Policy Accuracy Score (PAS) rubric implementation; (3) a 30-scenario NIST SP 800-53–derived dataset with reference OPA/Rego policies, reference Terraform configurations, and paired compliant / non-compliant OPA test fixtures; (4) the 180-run experiment that produced the results reported in the paper (3 commercial LLMs × 2 generation methods × 30 scenarios); (5) cached model outputs and OPA / Checkov proof artefacts enabling offline re-scoring without API access; (6) the paper-ready figures and tables generated from the canonical results CSV. The canonical results CSV has SHA-256 42f6ec840d7745653749d7a39aea6e338afbd81e30b8f5df7099b31e1de2643c.</p> |
| title | An Evaluation Framework for LLM-Driven Regulatory-to-Policy-as-Code Translation - Replication Package |
| topic | large language models evaluation framework DevSecOps policy as code regulatory compliance retrieval-augmented generation NIST SP 800-53 Open Policy Agent Rego Terraform |
| url | https://doi.org/10.5281/zenodo.20399040 |