Project Alexandria: Towards Freeing Scientific Knowledge from Copyright Burdens via LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Schuhmann, Christoph, Rabby, Gollam, Prabhu, Ameya, Ahmed, Tawsif, Hochlehnert, Andreas, Nguyen, Huu, Akinci, Nick, Schmidt, Ludwig, Kaczmarczyk, Robert, Auer, Sören, Jitsev, Jenia, Bethge, Matthias
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916695463952384
author Schuhmann, Christoph
Rabby, Gollam
Prabhu, Ameya
Ahmed, Tawsif
Hochlehnert, Andreas
Nguyen, Huu
Akinci, Nick
Schmidt, Ludwig
Kaczmarczyk, Robert
Auer, Sören
Jitsev, Jenia
Bethge, Matthias
author_facet Schuhmann, Christoph
Rabby, Gollam
Prabhu, Ameya
Ahmed, Tawsif
Hochlehnert, Andreas
Nguyen, Huu
Akinci, Nick
Schmidt, Ludwig
Kaczmarczyk, Robert
Auer, Sören
Jitsev, Jenia
Bethge, Matthias
contents Paywalls, licenses and copyright rules often restrict the broad dissemination and reuse of scientific knowledge. We take the position that it is both legally and technically feasible to extract the scientific knowledge in scholarly texts. Current methods, like text embeddings, fail to reliably preserve factual content, and simple paraphrasing may not be legally sound. We propose a new idea for the community to adopt: convert scholarly documents into knowledge preserving, but style agnostic representations we term Knowledge Units using LLMs. These units use structured data capturing entities, attributes and relationships without stylistic content. We provide evidence that Knowledge Units (1) form a legally defensible framework for sharing knowledge from copyrighted research texts, based on legal analyses of German copyright law and U.S. Fair Use doctrine, and (2) preserve most (~95\%) factual knowledge from original text, measured by MCQ performance on facts from the original copyrighted text across four research domains. Freeing scientific knowledge from copyright promises transformative benefits for scientific research and education by allowing language models to reuse important facts from copyrighted text. To support this, we share open-source tools for converting research documents into Knowledge Units. Overall, our work posits the feasibility of democratizing access to scientific knowledge while respecting copyright.
format Preprint
id arxiv_https___arxiv_org_abs_2502_19413
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Project Alexandria: Towards Freeing Scientific Knowledge from Copyright Burdens via LLMs
Schuhmann, Christoph
Rabby, Gollam
Prabhu, Ameya
Ahmed, Tawsif
Hochlehnert, Andreas
Nguyen, Huu
Akinci, Nick
Schmidt, Ludwig
Kaczmarczyk, Robert
Auer, Sören
Jitsev, Jenia
Bethge, Matthias
Machine Learning
Artificial Intelligence
Computation and Language
Paywalls, licenses and copyright rules often restrict the broad dissemination and reuse of scientific knowledge. We take the position that it is both legally and technically feasible to extract the scientific knowledge in scholarly texts. Current methods, like text embeddings, fail to reliably preserve factual content, and simple paraphrasing may not be legally sound. We propose a new idea for the community to adopt: convert scholarly documents into knowledge preserving, but style agnostic representations we term Knowledge Units using LLMs. These units use structured data capturing entities, attributes and relationships without stylistic content. We provide evidence that Knowledge Units (1) form a legally defensible framework for sharing knowledge from copyrighted research texts, based on legal analyses of German copyright law and U.S. Fair Use doctrine, and (2) preserve most (~95\%) factual knowledge from original text, measured by MCQ performance on facts from the original copyrighted text across four research domains. Freeing scientific knowledge from copyright promises transformative benefits for scientific research and education by allowing language models to reuse important facts from copyrighted text. To support this, we share open-source tools for converting research documents into Knowledge Units. Overall, our work posits the feasibility of democratizing access to scientific knowledge while respecting copyright.
title Project Alexandria: Towards Freeing Scientific Knowledge from Copyright Burdens via LLMs
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2502.19413