GraphBPE: Molecular Graphs Meet Byte-Pair Encoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Yuchen, Póczos, Barnabás
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909272027168768
author Shen, Yuchen
Póczos, Barnabás
author_facet Shen, Yuchen
Póczos, Barnabás
contents With the increasing attention to molecular machine learning, various innovations have been made in designing better models or proposing more comprehensive benchmarks. However, less is studied on the data preprocessing schedule for molecular graphs, where a different view of the molecular graph could potentially boost the model's performance. Inspired by the Byte-Pair Encoding (BPE) algorithm, a subword tokenization method popularly adopted in Natural Language Processing, we propose GraphBPE, which tokenizes a molecular graph into different substructures and acts as a preprocessing schedule independent of the model architectures. Our experiments on 3 graph-level classification and 3 graph-level regression datasets show that data preprocessing could boost the performance of models for molecular graphs, and GraphBPE is effective for small classification datasets and it performs on par with other tokenization methods across different model architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2407_19039
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GraphBPE: Molecular Graphs Meet Byte-Pair Encoding
Shen, Yuchen
Póczos, Barnabás
Machine Learning
Artificial Intelligence
Chemical Physics
Biomolecules
With the increasing attention to molecular machine learning, various innovations have been made in designing better models or proposing more comprehensive benchmarks. However, less is studied on the data preprocessing schedule for molecular graphs, where a different view of the molecular graph could potentially boost the model's performance. Inspired by the Byte-Pair Encoding (BPE) algorithm, a subword tokenization method popularly adopted in Natural Language Processing, we propose GraphBPE, which tokenizes a molecular graph into different substructures and acts as a preprocessing schedule independent of the model architectures. Our experiments on 3 graph-level classification and 3 graph-level regression datasets show that data preprocessing could boost the performance of models for molecular graphs, and GraphBPE is effective for small classification datasets and it performs on par with other tokenization methods across different model architectures.
title GraphBPE: Molecular Graphs Meet Byte-Pair Encoding
topic Machine Learning
Artificial Intelligence
Chemical Physics
Biomolecules
url https://arxiv.org/abs/2407.19039