MolPILE -- large-scale, diverse dataset for molecular representation learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Adamczyk, Jakub, Poziemski, Jakub, Job, Franciszek, Król, Mateusz, Makowski, Maciej
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909805720895488
author Adamczyk, Jakub
Poziemski, Jakub
Job, Franciszek
Król, Mateusz
Makowski, Maciej
author_facet Adamczyk, Jakub
Poziemski, Jakub
Job, Franciszek
Król, Mateusz
Makowski, Maciej
contents The size, diversity, and quality of pretraining datasets critically determine the generalization ability of foundation models. Despite their growing importance in chemoinformatics, the effectiveness of molecular representation learning has been hindered by limitations in existing small molecule datasets. To address this gap, we present MolPILE, large-scale, diverse, and rigorously curated collection of 222 million compounds, constructed from 6 large-scale databases using an automated curation pipeline. We present a comprehensive analysis of current pretraining datasets, highlighting considerable shortcomings for training ML models, and demonstrate how retraining existing models on MolPILE yields improvements in generalization performance. This work provides a standardized resource for model training, addressing the pressing need for an ImageNet-like dataset in molecular chemistry.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18353
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MolPILE -- large-scale, diverse dataset for molecular representation learning
Adamczyk, Jakub
Poziemski, Jakub
Job, Franciszek
Król, Mateusz
Makowski, Maciej
Machine Learning
The size, diversity, and quality of pretraining datasets critically determine the generalization ability of foundation models. Despite their growing importance in chemoinformatics, the effectiveness of molecular representation learning has been hindered by limitations in existing small molecule datasets. To address this gap, we present MolPILE, large-scale, diverse, and rigorously curated collection of 222 million compounds, constructed from 6 large-scale databases using an automated curation pipeline. We present a comprehensive analysis of current pretraining datasets, highlighting considerable shortcomings for training ML models, and demonstrate how retraining existing models on MolPILE yields improvements in generalization performance. This work provides a standardized resource for model training, addressing the pressing need for an ImageNet-like dataset in molecular chemistry.
title MolPILE -- large-scale, diverse dataset for molecular representation learning
topic Machine Learning
url https://arxiv.org/abs/2509.18353