Thinking like a CHEMIST: Combined Heterogeneous Embedding Model Integrating Structure and Tokens

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rekut, Nikolai, Orlov, Alexey, Ziu, Klea, Starykh, Elizaveta, Takac, Martin, Beznosikov, Aleksandr
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916756113588224
author Rekut, Nikolai
Orlov, Alexey
Ziu, Klea
Starykh, Elizaveta
Takac, Martin
Beznosikov, Aleksandr
author_facet Rekut, Nikolai
Orlov, Alexey
Ziu, Klea
Starykh, Elizaveta
Takac, Martin
Beznosikov, Aleksandr
contents Representing molecular structures effectively in chemistry remains a challenging task. Language models and graph-based models are extensively utilized within this domain, consistently achieving state-of-the-art results across an array of tasks. However, the prevailing practice of representing chemical compounds in the SMILES format - used by most data sets and many language models - presents notable limitations as a training data format. In this study, we present a novel approach that decomposes molecules into substructures and computes descriptor-based representations for these fragments, providing more detailed and chemically relevant input for model training. We use this substructure and descriptor data as input for language model and also propose a bimodal architecture that integrates this language model with graph-based models. As LM we use RoBERTa, Graph Isomorphism Networks (GIN), Graph Convolutional Networks (GCN) and Graphormer as graph ones. Our framework shows notable improvements over traditional methods in various tasks such as Quantitative Structure-Activity Relationship (QSAR) prediction.
format Preprint
id arxiv_https___arxiv_org_abs_2502_17986
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Thinking like a CHEMIST: Combined Heterogeneous Embedding Model Integrating Structure and Tokens
Rekut, Nikolai
Orlov, Alexey
Ziu, Klea
Starykh, Elizaveta
Takac, Martin
Beznosikov, Aleksandr
Machine Learning
Artificial Intelligence
Representing molecular structures effectively in chemistry remains a challenging task. Language models and graph-based models are extensively utilized within this domain, consistently achieving state-of-the-art results across an array of tasks. However, the prevailing practice of representing chemical compounds in the SMILES format - used by most data sets and many language models - presents notable limitations as a training data format. In this study, we present a novel approach that decomposes molecules into substructures and computes descriptor-based representations for these fragments, providing more detailed and chemically relevant input for model training. We use this substructure and descriptor data as input for language model and also propose a bimodal architecture that integrates this language model with graph-based models. As LM we use RoBERTa, Graph Isomorphism Networks (GIN), Graph Convolutional Networks (GCN) and Graphormer as graph ones. Our framework shows notable improvements over traditional methods in various tasks such as Quantitative Structure-Activity Relationship (QSAR) prediction.
title Thinking like a CHEMIST: Combined Heterogeneous Embedding Model Integrating Structure and Tokens
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2502.17986