HIGHT: Hierarchical Graph Tokenization for Molecule-Language Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yongqiang, Yao, Quanming, Zhang, Juzheng, Cheng, James, Bian, Yatao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908395623153664
author Chen, Yongqiang
Yao, Quanming
Zhang, Juzheng
Cheng, James
Bian, Yatao
author_facet Chen, Yongqiang
Yao, Quanming
Zhang, Juzheng
Cheng, James
Bian, Yatao
contents Recently, there has been a surge of interest in extending the success of large language models (LLMs) from texts to molecules. Most existing approaches adopt a graph neural network to represent a molecule as a series of node tokens for molecule-language alignment, which, however, have overlooked the inherent hierarchical structures in molecules. Notably, higher-order molecular structures contain rich semantics of functional groups, which encode crucial biochemical functionalities of the molecules. We show that neglecting the hierarchical information in tokenization will lead to subpar molecule-language alignment and severe hallucination. To address this limitation, we propose HIerarchical GrapH Tokenization (HIGHT). HIGHT employs a hierarchical graph tokenizer that encodes the hierarchy of atom, motif, and molecular levels of informative tokens to improve the molecular perception of LLMs. HIGHT also adopts an augmented instruction tuning dataset, enriched with the hierarchical graph information, to further enhance the molecule-language alignment. Extensive experiments on 14 real-world benchmarks verify the effectiveness of HIGHT in reducing hallucination by 40%, and significant improvements in various molecule-language downstream tasks. The project is available at https: //higraphllm.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2406_14021
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HIGHT: Hierarchical Graph Tokenization for Molecule-Language Alignment
Chen, Yongqiang
Yao, Quanming
Zhang, Juzheng
Cheng, James
Bian, Yatao
Computation and Language
Machine Learning
Quantitative Methods
Recently, there has been a surge of interest in extending the success of large language models (LLMs) from texts to molecules. Most existing approaches adopt a graph neural network to represent a molecule as a series of node tokens for molecule-language alignment, which, however, have overlooked the inherent hierarchical structures in molecules. Notably, higher-order molecular structures contain rich semantics of functional groups, which encode crucial biochemical functionalities of the molecules. We show that neglecting the hierarchical information in tokenization will lead to subpar molecule-language alignment and severe hallucination. To address this limitation, we propose HIerarchical GrapH Tokenization (HIGHT). HIGHT employs a hierarchical graph tokenizer that encodes the hierarchy of atom, motif, and molecular levels of informative tokens to improve the molecular perception of LLMs. HIGHT also adopts an augmented instruction tuning dataset, enriched with the hierarchical graph information, to further enhance the molecule-language alignment. Extensive experiments on 14 real-world benchmarks verify the effectiveness of HIGHT in reducing hallucination by 40%, and significant improvements in various molecule-language downstream tasks. The project is available at https: //higraphllm.github.io/.
title HIGHT: Hierarchical Graph Tokenization for Molecule-Language Alignment
topic Computation and Language
Machine Learning
Quantitative Methods
url https://arxiv.org/abs/2406.14021