Secure Retrieval-Augmented Generation: Preventing Data Leakage with Provenance and Policy Enforcement

Fuente: Zenodo
Saved in:
Bibliographic Details
Main Author: Narendra Bhargav Boggarapu
Format: Recurso digital
Published: Zenodo 2026
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866901556060749824
author Narendra Bhargav Boggarapu
author_facet Narendra Bhargav Boggarapu
contents <p>Retrieval-Augmented Generation (RAG) has become a paradigm architectural design to implement large language models in industries that face some form of regulation, but its probabilistic nature of retrieval creates a high risk of data leaks where access controls are not enforced or are used improperly. Secure RAG must have a comprehensive governance posture covering ingestion, indexing, retrieval, and generation mediated by identity-aware filtering, attribute-based access control, and policy-as-code frameworks that are version-controlled and auditably separate. Threats in this list are prompt injection, entitlement bypass, index poisoning, and generation-stage inferential disclosure, which need different prevention measures at the correct stage of the pipeline. The structural basis of the consistent enforcement is a two-plane architecture between the data plane and the control plane, where Open Policy Agent is the common decision-making and verifiable evidence bundle that meets the traceability considerations of the IEEE P7001 transparency standard. Assessment should be ongoing as opposed to periodic, and in terms of leakage rate, entitlement violation rate, provenance fidelity, and refusal correctness against a security harness that is consistent with the four core governance functions of the NIST AI RMF. The outcome is a deployment model where the enforcement of policy can be measured, provenance can be demonstrated to regulatory inquisitions, and leakage risk is structurally constrained as opposed to the widely held belief that this risk is mitigated.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_19070590
institution Zenodo
language
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Secure Retrieval-Augmented Generation: Preventing Data Leakage with Provenance and Policy Enforcement
Narendra Bhargav Boggarapu
<p>Retrieval-Augmented Generation (RAG) has become a paradigm architectural design to implement large language models in industries that face some form of regulation, but its probabilistic nature of retrieval creates a high risk of data leaks where access controls are not enforced or are used improperly. Secure RAG must have a comprehensive governance posture covering ingestion, indexing, retrieval, and generation mediated by identity-aware filtering, attribute-based access control, and policy-as-code frameworks that are version-controlled and auditably separate. Threats in this list are prompt injection, entitlement bypass, index poisoning, and generation-stage inferential disclosure, which need different prevention measures at the correct stage of the pipeline. The structural basis of the consistent enforcement is a two-plane architecture between the data plane and the control plane, where Open Policy Agent is the common decision-making and verifiable evidence bundle that meets the traceability considerations of the IEEE P7001 transparency standard. Assessment should be ongoing as opposed to periodic, and in terms of leakage rate, entitlement violation rate, provenance fidelity, and refusal correctness against a security harness that is consistent with the four core governance functions of the NIST AI RMF. The outcome is a deployment model where the enforcement of policy can be measured, provenance can be demonstrated to regulatory inquisitions, and leakage risk is structurally constrained as opposed to the widely held belief that this risk is mitigated.</p>
title Secure Retrieval-Augmented Generation: Preventing Data Leakage with Provenance and Policy Enforcement
url https://doi.org/10.5281/zenodo.19070590