Transparent NLP: Using RAG and LLM Alignment for Privacy Q&A

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Leschanowsky, Anna, Kolagar, Zahra, Çano, Erion, Habernal, Ivan, Hallinan, Dara, Habets, Emanuël A. P., Popp, Birgit
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917917658972160
author Leschanowsky, Anna
Kolagar, Zahra
Çano, Erion
Habernal, Ivan
Hallinan, Dara
Habets, Emanuël A. P.
Popp, Birgit
author_facet Leschanowsky, Anna
Kolagar, Zahra
Çano, Erion
Habernal, Ivan
Hallinan, Dara
Habets, Emanuël A. P.
Popp, Birgit
contents The transparency principle of the General Data Protection Regulation (GDPR) requires data processing information to be clear, precise, and accessible. While language models show promise in this context, their probabilistic nature complicates truthfulness and comprehensibility. This paper examines state-of-the-art Retrieval Augmented Generation (RAG) systems enhanced with alignment techniques to fulfill GDPR obligations. We evaluate RAG systems incorporating an alignment module like Rewindable Auto-regressive Inference (RAIN) and our proposed multidimensional extension, MultiRAIN, using a Privacy Q&A dataset. Responses are optimized for preciseness and comprehensibility and are assessed through 21 metrics, including deterministic and large language model-based evaluations. Our results show that RAG systems with an alignment module outperform baseline RAG systems on most metrics, though none fully match human answers. Principal component analysis of the results reveals complex interactions between metrics, highlighting the need to refine metrics. This study provides a foundation for integrating advanced natural language processing systems into legal compliance frameworks.
format Preprint
id arxiv_https___arxiv_org_abs_2502_06652
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Transparent NLP: Using RAG and LLM Alignment for Privacy Q&A
Leschanowsky, Anna
Kolagar, Zahra
Çano, Erion
Habernal, Ivan
Hallinan, Dara
Habets, Emanuël A. P.
Popp, Birgit
Computation and Language
The transparency principle of the General Data Protection Regulation (GDPR) requires data processing information to be clear, precise, and accessible. While language models show promise in this context, their probabilistic nature complicates truthfulness and comprehensibility. This paper examines state-of-the-art Retrieval Augmented Generation (RAG) systems enhanced with alignment techniques to fulfill GDPR obligations. We evaluate RAG systems incorporating an alignment module like Rewindable Auto-regressive Inference (RAIN) and our proposed multidimensional extension, MultiRAIN, using a Privacy Q&A dataset. Responses are optimized for preciseness and comprehensibility and are assessed through 21 metrics, including deterministic and large language model-based evaluations. Our results show that RAG systems with an alignment module outperform baseline RAG systems on most metrics, though none fully match human answers. Principal component analysis of the results reveals complex interactions between metrics, highlighting the need to refine metrics. This study provides a foundation for integrating advanced natural language processing systems into legal compliance frameworks.
title Transparent NLP: Using RAG and LLM Alignment for Privacy Q&A
topic Computation and Language
url https://arxiv.org/abs/2502.06652