Transparent NLP: Using RAG and LLM Alignment for Privacy Q&A
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866917917658972160 |
|---|---|
| author | Leschanowsky, Anna Kolagar, Zahra Çano, Erion Habernal, Ivan Hallinan, Dara Habets, Emanuël A. P. Popp, Birgit |
| author_facet | Leschanowsky, Anna Kolagar, Zahra Çano, Erion Habernal, Ivan Hallinan, Dara Habets, Emanuël A. P. Popp, Birgit |
| contents | The transparency principle of the General Data Protection Regulation (GDPR) requires data processing information to be clear, precise, and accessible. While language models show promise in this context, their probabilistic nature complicates truthfulness and comprehensibility.
This paper examines state-of-the-art Retrieval Augmented Generation (RAG) systems enhanced with alignment techniques to fulfill GDPR obligations. We evaluate RAG systems incorporating an alignment module like Rewindable Auto-regressive Inference (RAIN) and our proposed multidimensional extension, MultiRAIN, using a Privacy Q&A dataset. Responses are optimized for preciseness and comprehensibility and are assessed through 21 metrics, including deterministic and large language model-based evaluations.
Our results show that RAG systems with an alignment module outperform baseline RAG systems on most metrics, though none fully match human answers. Principal component analysis of the results reveals complex interactions between metrics, highlighting the need to refine metrics. This study provides a foundation for integrating advanced natural language processing systems into legal compliance frameworks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_06652 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Transparent NLP: Using RAG and LLM Alignment for Privacy Q&A Leschanowsky, Anna Kolagar, Zahra Çano, Erion Habernal, Ivan Hallinan, Dara Habets, Emanuël A. P. Popp, Birgit Computation and Language The transparency principle of the General Data Protection Regulation (GDPR) requires data processing information to be clear, precise, and accessible. While language models show promise in this context, their probabilistic nature complicates truthfulness and comprehensibility. This paper examines state-of-the-art Retrieval Augmented Generation (RAG) systems enhanced with alignment techniques to fulfill GDPR obligations. We evaluate RAG systems incorporating an alignment module like Rewindable Auto-regressive Inference (RAIN) and our proposed multidimensional extension, MultiRAIN, using a Privacy Q&A dataset. Responses are optimized for preciseness and comprehensibility and are assessed through 21 metrics, including deterministic and large language model-based evaluations. Our results show that RAG systems with an alignment module outperform baseline RAG systems on most metrics, though none fully match human answers. Principal component analysis of the results reveals complex interactions between metrics, highlighting the need to refine metrics. This study provides a foundation for integrating advanced natural language processing systems into legal compliance frameworks. |
| title | Transparent NLP: Using RAG and LLM Alignment for Privacy Q&A |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2502.06652 |