Entailment-Driven Privacy Policy Classification with LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Silva, Bhanuka, Denipitiyage, Dishanika, Seneviratne, Suranga, Mahanti, Anirban, Seneviratne, Aruna
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929515019632640
author Silva, Bhanuka
Denipitiyage, Dishanika
Seneviratne, Suranga
Mahanti, Anirban
Seneviratne, Aruna
author_facet Silva, Bhanuka
Denipitiyage, Dishanika
Seneviratne, Suranga
Mahanti, Anirban
Seneviratne, Aruna
contents While many online services provide privacy policies for end users to read and understand what personal data are being collected, these documents are often lengthy and complicated. As a result, the vast majority of users do not read them at all, leading to data collection under uninformed consent. Several attempts have been made to make privacy policies more user friendly by summarising them, providing automatic annotations or labels for key sections, or by offering chat interfaces to ask specific questions. With recent advances in Large Language Models (LLMs), there is an opportunity to develop more effective tools to parse privacy policies and help users make informed decisions. In this paper, we propose an entailment-driven LLM based framework to classify paragraphs of privacy policies into meaningful labels that are easily understood by users. The results demonstrate that our framework outperforms traditional LLM methods, improving the F1 score in average by 11.2%. Additionally, our framework provides inherently explainable and meaningful predictions.
format Preprint
id arxiv_https___arxiv_org_abs_2409_16621
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Entailment-Driven Privacy Policy Classification with LLMs
Silva, Bhanuka
Denipitiyage, Dishanika
Seneviratne, Suranga
Mahanti, Anirban
Seneviratne, Aruna
Artificial Intelligence
While many online services provide privacy policies for end users to read and understand what personal data are being collected, these documents are often lengthy and complicated. As a result, the vast majority of users do not read them at all, leading to data collection under uninformed consent. Several attempts have been made to make privacy policies more user friendly by summarising them, providing automatic annotations or labels for key sections, or by offering chat interfaces to ask specific questions. With recent advances in Large Language Models (LLMs), there is an opportunity to develop more effective tools to parse privacy policies and help users make informed decisions. In this paper, we propose an entailment-driven LLM based framework to classify paragraphs of privacy policies into meaningful labels that are easily understood by users. The results demonstrate that our framework outperforms traditional LLM methods, improving the F1 score in average by 11.2%. Additionally, our framework provides inherently explainable and meaningful predictions.
title Entailment-Driven Privacy Policy Classification with LLMs
topic Artificial Intelligence
url https://arxiv.org/abs/2409.16621