A Multi-lingual Dataset of Classified Paragraphs from Open Access Scientific Publications

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autore principale: Jeangirard, Eric
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911230928617472
author Jeangirard, Eric
author_facet Jeangirard, Eric
contents We present a dataset of 833k paragraphs extracted from CC-BY licensed scientific publications, classified into four categories: acknowledgments, data mentions, software/code mentions, and clinical trial mentions. The paragraphs are primarily in English and French, with additional European languages represented. Each paragraph is annotated with language identification (using fastText) and scientific domain (from OpenAlex). This dataset, derived from the French Open Science Monitor corpus and processed using GROBID, enables training of text classification models and development of named entity recognition systems for scientific literature mining. The dataset is publicly available on HuggingFace https://doi.org/10.57967/hf/6679 under a CC-BY license.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21762
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Multi-lingual Dataset of Classified Paragraphs from Open Access Scientific Publications
Jeangirard, Eric
Computation and Language
Digital Libraries
We present a dataset of 833k paragraphs extracted from CC-BY licensed scientific publications, classified into four categories: acknowledgments, data mentions, software/code mentions, and clinical trial mentions. The paragraphs are primarily in English and French, with additional European languages represented. Each paragraph is annotated with language identification (using fastText) and scientific domain (from OpenAlex). This dataset, derived from the French Open Science Monitor corpus and processed using GROBID, enables training of text classification models and development of named entity recognition systems for scientific literature mining. The dataset is publicly available on HuggingFace https://doi.org/10.57967/hf/6679 under a CC-BY license.
title A Multi-lingual Dataset of Classified Paragraphs from Open Access Scientific Publications
topic Computation and Language
Digital Libraries
url https://arxiv.org/abs/2510.21762