Two for the Price of One: Integrating Large Language Models to Learn Biophysical Interactions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Clark, Joseph D., Dean, Tanner J., Shukla, Diwakar
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916663510695936
author Clark, Joseph D.
Dean, Tanner J.
Shukla, Diwakar
author_facet Clark, Joseph D.
Dean, Tanner J.
Shukla, Diwakar
contents Deep learning models have become fundamental tools in drug design. In particular, large language models trained on biochemical sequences learn feature vectors that guide drug discovery through virtual screening. However, such models do not capture the molecular interactions important for binding affinity and specificity. Therefore, there is a need to 'compose' representations from distinct biological modalities to effectively represent molecular complexes. We present an overview of the methods to combine molecular representations and propose that future work should balance computational efficiency and expressiveness. Specifically, we argue that improvements in both speed and accuracy are possible by learning to merge the representations from internal layers of domain specific biological language models. We demonstrate that 'composing' biochemical language models performs similar or better than standard methods representing molecular interactions despite having significantly fewer features. Finally, we discuss recent methods for interpreting and democratizing large language models that could aid the development of interaction aware foundation models for biology, as well as their shortcomings.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21017
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Two for the Price of One: Integrating Large Language Models to Learn Biophysical Interactions
Clark, Joseph D.
Dean, Tanner J.
Shukla, Diwakar
Biomolecules
Quantitative Methods
Deep learning models have become fundamental tools in drug design. In particular, large language models trained on biochemical sequences learn feature vectors that guide drug discovery through virtual screening. However, such models do not capture the molecular interactions important for binding affinity and specificity. Therefore, there is a need to 'compose' representations from distinct biological modalities to effectively represent molecular complexes. We present an overview of the methods to combine molecular representations and propose that future work should balance computational efficiency and expressiveness. Specifically, we argue that improvements in both speed and accuracy are possible by learning to merge the representations from internal layers of domain specific biological language models. We demonstrate that 'composing' biochemical language models performs similar or better than standard methods representing molecular interactions despite having significantly fewer features. Finally, we discuss recent methods for interpreting and democratizing large language models that could aid the development of interaction aware foundation models for biology, as well as their shortcomings.
title Two for the Price of One: Integrating Large Language Models to Learn Biophysical Interactions
topic Biomolecules
Quantitative Methods
url https://arxiv.org/abs/2503.21017