A Foundation Chemical Language Model for Comprehensive Fragment-Based Drug Discovery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ho, Alexander, Lee, Sukyeong, Tsai, Francis T. F.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908556466323456
author Ho, Alexander
Lee, Sukyeong
Tsai, Francis T. F.
author_facet Ho, Alexander
Lee, Sukyeong
Tsai, Francis T. F.
contents We introduce FragAtlas-62M, a specialized foundation model trained on the largest fragment dataset to date. Built on the complete ZINC-22 fragment subset comprising over 62 million molecules, it achieves unprecedented coverage of fragment chemical space. Our GPT-2 based model (42.7M parameters) generates 99.90% chemically valid fragments. Validation across 12 descriptors and three fingerprint methods shows generated fragments closely match the training distribution (all effect sizes < 0.4). The model retains 53.6% of known ZINC fragments while producing 22% novel structures with practical relevance. We release FragAtlas-62M with training code, preprocessed data, documentation, and model weights to accelerate adoption.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19586
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Foundation Chemical Language Model for Comprehensive Fragment-Based Drug Discovery
Ho, Alexander
Lee, Sukyeong
Tsai, Francis T. F.
Machine Learning
Artificial Intelligence
Biomolecules
We introduce FragAtlas-62M, a specialized foundation model trained on the largest fragment dataset to date. Built on the complete ZINC-22 fragment subset comprising over 62 million molecules, it achieves unprecedented coverage of fragment chemical space. Our GPT-2 based model (42.7M parameters) generates 99.90% chemically valid fragments. Validation across 12 descriptors and three fingerprint methods shows generated fragments closely match the training distribution (all effect sizes < 0.4). The model retains 53.6% of known ZINC fragments while producing 22% novel structures with practical relevance. We release FragAtlas-62M with training code, preprocessed data, documentation, and model weights to accelerate adoption.
title A Foundation Chemical Language Model for Comprehensive Fragment-Based Drug Discovery
topic Machine Learning
Artificial Intelligence
Biomolecules
url https://arxiv.org/abs/2509.19586