MolTextNet: A Two-Million Molecule-Text Dataset for Multimodal Molecular Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhu, Yihan, Liu, Gang, Inae, Eric, Jiang, Meng
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915315344998400
author Zhu, Yihan
Liu, Gang
Inae, Eric
Jiang, Meng
author_facet Zhu, Yihan
Liu, Gang
Inae, Eric
Jiang, Meng
contents Small molecules are essential to drug discovery, and graph-language models hold promise for learning molecular properties and functions from text. However, existing molecule-text datasets are limited in scale and informativeness, restricting the training of generalizable multimodal models. We present MolTextNet, a dataset of 2.5 million high-quality molecule-text pairs designed to overcome these limitations. To construct it, we propose a synthetic text generation pipeline that integrates structural features, computed properties, bioactivity data, and synthetic complexity. Using GPT-4o-mini, we create structured descriptions for 2.5 million molecules from ChEMBL35, with text over 10 times longer than prior datasets. MolTextNet supports diverse downstream tasks, including property prediction and structure retrieval. Pretraining CLIP-style models with Graph Neural Networks and ModernBERT on MolTextNet yields improved performance, highlighting its potential for advancing foundational multimodal modeling in molecular science. Our dataset is available at https://huggingface.co/datasets/liuganghuggingface/moltextnet.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00009
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MolTextNet: A Two-Million Molecule-Text Dataset for Multimodal Molecular Learning
Zhu, Yihan
Liu, Gang
Inae, Eric
Jiang, Meng
Biomolecules
Artificial Intelligence
Small molecules are essential to drug discovery, and graph-language models hold promise for learning molecular properties and functions from text. However, existing molecule-text datasets are limited in scale and informativeness, restricting the training of generalizable multimodal models. We present MolTextNet, a dataset of 2.5 million high-quality molecule-text pairs designed to overcome these limitations. To construct it, we propose a synthetic text generation pipeline that integrates structural features, computed properties, bioactivity data, and synthetic complexity. Using GPT-4o-mini, we create structured descriptions for 2.5 million molecules from ChEMBL35, with text over 10 times longer than prior datasets. MolTextNet supports diverse downstream tasks, including property prediction and structure retrieval. Pretraining CLIP-style models with Graph Neural Networks and ModernBERT on MolTextNet yields improved performance, highlighting its potential for advancing foundational multimodal modeling in molecular science. Our dataset is available at https://huggingface.co/datasets/liuganghuggingface/moltextnet.
title MolTextNet: A Two-Million Molecule-Text Dataset for Multimodal Molecular Learning
topic Biomolecules
Artificial Intelligence
url https://arxiv.org/abs/2506.00009