What Ails Generative Structure-based Drug Design: Expressivity is Too Little or Too Much?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Karczewski, Rafał, Kaski, Samuel, Heinonen, Markus, Garg, Vikas
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917943566139392
author Karczewski, Rafał
Kaski, Samuel
Heinonen, Markus
Garg, Vikas
author_facet Karczewski, Rafał
Kaski, Samuel
Heinonen, Markus
Garg, Vikas
contents Several generative models with elaborate training and sampling procedures have been proposed to accelerate structure-based drug design (SBDD); however, their empirical performance turns out to be suboptimal. We seek to better understand this phenomenon from both theoretical and empirical perspectives. Since most of these models apply graph neural networks (GNNs), one may suspect that they inherit the representational limitations of GNNs. We analyze this aspect, establishing the first such results for protein-ligand complexes. A plausible counterview may attribute the underperformance of these models to their excessive parameterizations, inducing expressivity at the expense of generalization. We investigate this possibility with a simple metric-aware approach that learns an economical surrogate for affinity to infer an unlabelled molecular graph and optimizes for labels conditioned on this graph and molecular properties. The resulting model achieves state-of-the-art results using 100x fewer trainable parameters and affords up to 1000x speedup. Collectively, our findings underscore the need to reassess and redirect the existing paradigm and efforts for SBDD. Code is available at https://github.com/rafalkarczewski/SimpleSBDD.
format Preprint
id arxiv_https___arxiv_org_abs_2408_06050
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle What Ails Generative Structure-based Drug Design: Expressivity is Too Little or Too Much?
Karczewski, Rafał
Kaski, Samuel
Heinonen, Markus
Garg, Vikas
Machine Learning
Biomolecules
Several generative models with elaborate training and sampling procedures have been proposed to accelerate structure-based drug design (SBDD); however, their empirical performance turns out to be suboptimal. We seek to better understand this phenomenon from both theoretical and empirical perspectives. Since most of these models apply graph neural networks (GNNs), one may suspect that they inherit the representational limitations of GNNs. We analyze this aspect, establishing the first such results for protein-ligand complexes. A plausible counterview may attribute the underperformance of these models to their excessive parameterizations, inducing expressivity at the expense of generalization. We investigate this possibility with a simple metric-aware approach that learns an economical surrogate for affinity to infer an unlabelled molecular graph and optimizes for labels conditioned on this graph and molecular properties. The resulting model achieves state-of-the-art results using 100x fewer trainable parameters and affords up to 1000x speedup. Collectively, our findings underscore the need to reassess and redirect the existing paradigm and efforts for SBDD. Code is available at https://github.com/rafalkarczewski/SimpleSBDD.
title What Ails Generative Structure-based Drug Design: Expressivity is Too Little or Too Much?
topic Machine Learning
Biomolecules
url https://arxiv.org/abs/2408.06050