MolReFlect: Towards In-Context Fine-grained Alignments between Molecules and Texts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jiatong, Liu, Yunqing, Liu, Wei, Le, Jingdi, Zhang, Di, Fan, Wenqi, Zhou, Dongzhan, Li, Yuqiang, Li, Qing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910173468033024
author Li, Jiatong
Liu, Yunqing
Liu, Wei
Le, Jingdi
Zhang, Di
Fan, Wenqi
Zhou, Dongzhan
Li, Yuqiang
Li, Qing
author_facet Li, Jiatong
Liu, Yunqing
Liu, Wei
Le, Jingdi
Zhang, Di
Fan, Wenqi
Zhou, Dongzhan
Li, Yuqiang
Li, Qing
contents Molecule discovery is a pivotal research field, impacting everything from medicine to materials. Recently, Large Language Models (LLMs) have been widely adopted in molecular understanding and generation, serving as a bridge between the molecular space and the natural language space, yet the alignment between molecules and their corresponding captions remains a significant challenge. Previous endeavors typically treat molecules as monolithic inputs, lacking an intermediate reasoning process and sacrificing explainability. In this work, we define fine-grained alignments as the precise correspondence between a molecule's sub-structures and the textual phrases that explain their properties. These alignments are crucial for LLMs to understand molecules in a more accurate and explainable manner. Normally, such fine-grained alignments require expert annotation, which is both costly and time-consuming. To allow LLMs to automatically label and learn the fine-grained alignments, we propose MolReFlect, a novel teacher-student framework, where a teacher LLM first generates and refines mappings between caption phrases and SMILES substructures and then explicitly teaches these detailed alignments to a student LLM. Experimental results demonstrate that MolReFlect enables LLMs to significantly outperform previous baselines, achieving the state-of-the-art performance in the molecule-caption translation task. Our codes are available via: https://github.com/phenixace/MolReFlect.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14721
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MolReFlect: Towards In-Context Fine-grained Alignments between Molecules and Texts
Li, Jiatong
Liu, Yunqing
Liu, Wei
Le, Jingdi
Zhang, Di
Fan, Wenqi
Zhou, Dongzhan
Li, Yuqiang
Li, Qing
Computation and Language
Machine Learning
Quantitative Methods
Molecule discovery is a pivotal research field, impacting everything from medicine to materials. Recently, Large Language Models (LLMs) have been widely adopted in molecular understanding and generation, serving as a bridge between the molecular space and the natural language space, yet the alignment between molecules and their corresponding captions remains a significant challenge. Previous endeavors typically treat molecules as monolithic inputs, lacking an intermediate reasoning process and sacrificing explainability. In this work, we define fine-grained alignments as the precise correspondence between a molecule's sub-structures and the textual phrases that explain their properties. These alignments are crucial for LLMs to understand molecules in a more accurate and explainable manner. Normally, such fine-grained alignments require expert annotation, which is both costly and time-consuming. To allow LLMs to automatically label and learn the fine-grained alignments, we propose MolReFlect, a novel teacher-student framework, where a teacher LLM first generates and refines mappings between caption phrases and SMILES substructures and then explicitly teaches these detailed alignments to a student LLM. Experimental results demonstrate that MolReFlect enables LLMs to significantly outperform previous baselines, achieving the state-of-the-art performance in the molecule-caption translation task. Our codes are available via: https://github.com/phenixace/MolReFlect.
title MolReFlect: Towards In-Context Fine-grained Alignments between Molecules and Texts
topic Computation and Language
Machine Learning
Quantitative Methods
url https://arxiv.org/abs/2411.14721