Overcoming Copyright Barriers in Corpus Distribution Through Non-Reversible Hashing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Amalvy, Arthur, Labatut, Vincent, Bost, Xavier, Huang, Hen-Hsen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910166331424768
author Amalvy, Arthur
Labatut, Vincent
Bost, Xavier
Huang, Hen-Hsen
author_facet Amalvy, Arthur
Labatut, Vincent
Bost, Xavier
Huang, Hen-Hsen
contents While annotated corpora are crucial in the field of natural language processing (NLP), those containing copyrighted material are difficult to exchange among researchers. Yet, such corpora are necessary to fully represent the diversity of data found in the wild in the context of NLP tasks. We tackle this issue by proposing a method to lawfully and publicly share the annotations of copyrighted literary texts. The corpus creator shares the annotations in clear, along with a non-reversible hashed version of the source material. The corpus user must own the source material, and apply the same hash function to their own tokens, in order to match them to the shared annotations. Crucially, our method is robust to reasonable divergences in the version of the copyrighted data owned by the user. As an illustration, we present alignment experiments on different editions of novels. Our results show that our method is able to correctly align 98.7 to 99.79% of tokens depending on the novel, provided the user version is sufficiently close to the corpus creator's version. We publicly release novelshare, a Python implementation of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2604_23412
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Overcoming Copyright Barriers in Corpus Distribution Through Non-Reversible Hashing
Amalvy, Arthur
Labatut, Vincent
Bost, Xavier
Huang, Hen-Hsen
Computation and Language
While annotated corpora are crucial in the field of natural language processing (NLP), those containing copyrighted material are difficult to exchange among researchers. Yet, such corpora are necessary to fully represent the diversity of data found in the wild in the context of NLP tasks. We tackle this issue by proposing a method to lawfully and publicly share the annotations of copyrighted literary texts. The corpus creator shares the annotations in clear, along with a non-reversible hashed version of the source material. The corpus user must own the source material, and apply the same hash function to their own tokens, in order to match them to the shared annotations. Crucially, our method is robust to reasonable divergences in the version of the copyrighted data owned by the user. As an illustration, we present alignment experiments on different editions of novels. Our results show that our method is able to correctly align 98.7 to 99.79% of tokens depending on the novel, provided the user version is sufficiently close to the corpus creator's version. We publicly release novelshare, a Python implementation of our method.
title Overcoming Copyright Barriers in Corpus Distribution Through Non-Reversible Hashing
topic Computation and Language
url https://arxiv.org/abs/2604.23412