NovoMolGen: Rethinking Molecular Language Model Pretraining

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chitsaz, Kamran, Balaji, Roshan, Fournier, Quentin, Bhatt, Nirav Pravinbhai, Chandar, Sarath
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916911786229760
author Chitsaz, Kamran
Balaji, Roshan
Fournier, Quentin
Bhatt, Nirav Pravinbhai
Chandar, Sarath
author_facet Chitsaz, Kamran
Balaji, Roshan
Fournier, Quentin
Bhatt, Nirav Pravinbhai
Chandar, Sarath
contents Designing de-novo molecules with desired property profiles requires efficient exploration of the vast chemical space ranging from $10^{23}$ to $10^{60}$ possible synthesizable candidates. While various deep generative models have been developed to design small molecules using diverse input representations, Molecular Large Language Models (Mol-LLMs) based on string representations have emerged as a scalable approach capable of exploring billions of molecules. However, there remains limited understanding regarding how standard language modeling practices such as textual representations, tokenization strategies, model size, and dataset scale impact molecular generation performance. In this work, we systematically investigate these critical aspects by introducing NovoMolGen, a family of transformer-based foundation models pretrained on 1.5 billion molecules for de-novo molecule generation. Through extensive empirical analyses, we identify a weak correlation between performance metrics measured during pretraining and actual downstream performance, revealing important distinctions between molecular and general NLP training dynamics. NovoMolGen establishes new state-of-the-art results, substantially outperforming prior Mol-LLMs and specialized generative models in both unconstrained and goal-directed molecular generation tasks, thus providing a robust foundation for advancing efficient and effective molecular modeling strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2508_13408
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NovoMolGen: Rethinking Molecular Language Model Pretraining
Chitsaz, Kamran
Balaji, Roshan
Fournier, Quentin
Bhatt, Nirav Pravinbhai
Chandar, Sarath
Machine Learning
Designing de-novo molecules with desired property profiles requires efficient exploration of the vast chemical space ranging from $10^{23}$ to $10^{60}$ possible synthesizable candidates. While various deep generative models have been developed to design small molecules using diverse input representations, Molecular Large Language Models (Mol-LLMs) based on string representations have emerged as a scalable approach capable of exploring billions of molecules. However, there remains limited understanding regarding how standard language modeling practices such as textual representations, tokenization strategies, model size, and dataset scale impact molecular generation performance. In this work, we systematically investigate these critical aspects by introducing NovoMolGen, a family of transformer-based foundation models pretrained on 1.5 billion molecules for de-novo molecule generation. Through extensive empirical analyses, we identify a weak correlation between performance metrics measured during pretraining and actual downstream performance, revealing important distinctions between molecular and general NLP training dynamics. NovoMolGen establishes new state-of-the-art results, substantially outperforming prior Mol-LLMs and specialized generative models in both unconstrained and goal-directed molecular generation tasks, thus providing a robust foundation for advancing efficient and effective molecular modeling strategies.
title NovoMolGen: Rethinking Molecular Language Model Pretraining
topic Machine Learning
url https://arxiv.org/abs/2508.13408