Scaling Structure Aware Virtual Screening to Billions of Molecules with SPRINT

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: McNutt, Andrew T., Adduri, Abhinav K., Ellington, Caleb N., Dayao, Monica T., Xing, Eric P., Mohimani, Hosein, Koes, David R.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916574070308864
author McNutt, Andrew T.
Adduri, Abhinav K.
Ellington, Caleb N.
Dayao, Monica T.
Xing, Eric P.
Mohimani, Hosein
Koes, David R.
author_facet McNutt, Andrew T.
Adduri, Abhinav K.
Ellington, Caleb N.
Dayao, Monica T.
Xing, Eric P.
Mohimani, Hosein
Koes, David R.
contents Virtual screening of small molecules against protein targets can accelerate drug discovery and development by predicting drug-target interactions (DTIs). However, structure-based methods like molecular docking are too slow to allow for broad proteome-scale screens, limiting their application in screening for off-target effects or new molecular mechanisms. Recently, vector-based methods using protein language models (PLMs) have emerged as a complementary approach that bypasses explicit 3D structure modeling. Here, we develop SPRINT, a vector-based approach for screening entire chemical libraries against whole proteomes for DTIs and novel mechanisms of action. SPRINT improves on prior work by using a self-attention based architecture and structure-aware PLMs to learn drug-target co-embeddings for binder prediction, search, and retrieval. SPRINT achieves SOTA enrichment factors in virtual screening on LIT-PCBA, DTI classification benchmarks, and binding affinity prediction benchmarks, while providing interpretability in the form of residue-level attention maps. In addition to being both accurate and interpretable, SPRINT is ultra-fast: querying the whole human proteome against the ENAMINE Real Database (6.7B drugs) for the 100 most likely binders per protein takes 16 minutes. SPRINT promises to enable virtual screening at an unprecedented scale, opening up new opportunities for in silico drug repurposing and development. SPRINT is available on the web as ColabScreen: https://bit.ly/colab-screen
format Preprint
id arxiv_https___arxiv_org_abs_2411_15418
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scaling Structure Aware Virtual Screening to Billions of Molecules with SPRINT
McNutt, Andrew T.
Adduri, Abhinav K.
Ellington, Caleb N.
Dayao, Monica T.
Xing, Eric P.
Mohimani, Hosein
Koes, David R.
Biomolecules
Machine Learning
Virtual screening of small molecules against protein targets can accelerate drug discovery and development by predicting drug-target interactions (DTIs). However, structure-based methods like molecular docking are too slow to allow for broad proteome-scale screens, limiting their application in screening for off-target effects or new molecular mechanisms. Recently, vector-based methods using protein language models (PLMs) have emerged as a complementary approach that bypasses explicit 3D structure modeling. Here, we develop SPRINT, a vector-based approach for screening entire chemical libraries against whole proteomes for DTIs and novel mechanisms of action. SPRINT improves on prior work by using a self-attention based architecture and structure-aware PLMs to learn drug-target co-embeddings for binder prediction, search, and retrieval. SPRINT achieves SOTA enrichment factors in virtual screening on LIT-PCBA, DTI classification benchmarks, and binding affinity prediction benchmarks, while providing interpretability in the form of residue-level attention maps. In addition to being both accurate and interpretable, SPRINT is ultra-fast: querying the whole human proteome against the ENAMINE Real Database (6.7B drugs) for the 100 most likely binders per protein takes 16 minutes. SPRINT promises to enable virtual screening at an unprecedented scale, opening up new opportunities for in silico drug repurposing and development. SPRINT is available on the web as ColabScreen: https://bit.ly/colab-screen
title Scaling Structure Aware Virtual Screening to Billions of Molecules with SPRINT
topic Biomolecules
Machine Learning
url https://arxiv.org/abs/2411.15418