Privacy at Scale: Introducing the PrivaSeer Corpus of Web Privacy Policies

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Srinath, Mukund, Wilson, Shomir, Giles, C. Lee
Format: Preprint
Veröffentlicht: 2020
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910391020290048
author Srinath, Mukund
Wilson, Shomir
Giles, C. Lee
author_facet Srinath, Mukund
Wilson, Shomir
Giles, C. Lee
contents Organisations disclose their privacy practices by posting privacy policies on their website. Even though users often care about their digital privacy, they often don't read privacy policies since they require a significant investment in time and effort. Although natural language processing can help in privacy policy understanding, there has been a lack of large scale privacy policy corpora that could be used to analyse, understand, and simplify privacy policies. Thus, we create PrivaSeer, a corpus of over one million English language website privacy policies, which is significantly larger than any previously available corpus. We design a corpus creation pipeline which consists of crawling the web followed by filtering documents using language detection, document classification, duplicate and near-duplication removal, and content extraction. We investigate the composition of the corpus and show results from readability tests, document similarity, keyphrase extraction, and explored the corpus through topic modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2004_11131
institution arXiv
publishDate 2020
record_format arxiv
spellingShingle Privacy at Scale: Introducing the PrivaSeer Corpus of Web Privacy Policies
Srinath, Mukund
Wilson, Shomir
Giles, C. Lee
Information Retrieval
Cryptography and Security
Organisations disclose their privacy practices by posting privacy policies on their website. Even though users often care about their digital privacy, they often don't read privacy policies since they require a significant investment in time and effort. Although natural language processing can help in privacy policy understanding, there has been a lack of large scale privacy policy corpora that could be used to analyse, understand, and simplify privacy policies. Thus, we create PrivaSeer, a corpus of over one million English language website privacy policies, which is significantly larger than any previously available corpus. We design a corpus creation pipeline which consists of crawling the web followed by filtering documents using language detection, document classification, duplicate and near-duplication removal, and content extraction. We investigate the composition of the corpus and show results from readability tests, document similarity, keyphrase extraction, and explored the corpus through topic modeling.
title Privacy at Scale: Introducing the PrivaSeer Corpus of Web Privacy Policies
topic Information Retrieval
Cryptography and Security
url https://arxiv.org/abs/2004.11131