Tokenization for Molecular Foundation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wadell, Alexius, Bhutani, Anoushka, Viswanathan, Venkatasubramanian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914286849228800
author Wadell, Alexius
Bhutani, Anoushka
Viswanathan, Venkatasubramanian
author_facet Wadell, Alexius
Bhutani, Anoushka
Viswanathan, Venkatasubramanian
contents Text-based foundation models have become an important part of scientific discovery, with molecular foundation models accelerating advancements in material science and molecular design.However, existing models are constrained by closed-vocabulary tokenizers that capture only a fraction of molecular space. In this work, we systematically evaluate 34 tokenizers, including 19 chemistry-specific ones, and reveal significant gaps in their coverage of the SMILES molecular representation. To assess the impact of tokenizer choice, we introduce n-gram language models as a low-cost proxy and validate their effectiveness by pretraining and finetuning 18 RoBERTa-style encoders for molecular property prediction. To overcome the limitations of existing tokenizers, we propose two new tokenizers -- Smirk and Smirk-GPE -- with full coverage of the OpenSMILES specification. The proposed tokenizers systematically integrate nuclear, electronic, and geometric degrees of freedom; facilitating applications in pharmacology, agriculture, biology, and energy storage. Our results highlight the need for open-vocabulary modeling and chemically diverse benchmarks in cheminformatics.
format Preprint
id arxiv_https___arxiv_org_abs_2409_15370
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Tokenization for Molecular Foundation Models
Wadell, Alexius
Bhutani, Anoushka
Viswanathan, Venkatasubramanian
Machine Learning
Artificial Intelligence
Chemical Physics
Biomolecules
Text-based foundation models have become an important part of scientific discovery, with molecular foundation models accelerating advancements in material science and molecular design.However, existing models are constrained by closed-vocabulary tokenizers that capture only a fraction of molecular space. In this work, we systematically evaluate 34 tokenizers, including 19 chemistry-specific ones, and reveal significant gaps in their coverage of the SMILES molecular representation. To assess the impact of tokenizer choice, we introduce n-gram language models as a low-cost proxy and validate their effectiveness by pretraining and finetuning 18 RoBERTa-style encoders for molecular property prediction. To overcome the limitations of existing tokenizers, we propose two new tokenizers -- Smirk and Smirk-GPE -- with full coverage of the OpenSMILES specification. The proposed tokenizers systematically integrate nuclear, electronic, and geometric degrees of freedom; facilitating applications in pharmacology, agriculture, biology, and energy storage. Our results highlight the need for open-vocabulary modeling and chemically diverse benchmarks in cheminformatics.
title Tokenization for Molecular Foundation Models
topic Machine Learning
Artificial Intelligence
Chemical Physics
Biomolecules
url https://arxiv.org/abs/2409.15370