Navigating Chemical-Linguistic Sharing Space with Heterogeneous Molecular Encoding

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lv, Liuzhenghao, Li, Hao, Wang, Yu, Yan, Zhiyuan, Chen, Zijun, Lin, Zongying, Yuan, Li, Tian, Yonghong
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910773420228608
author Lv, Liuzhenghao
Li, Hao
Wang, Yu
Yan, Zhiyuan
Chen, Zijun
Lin, Zongying
Yuan, Li
Tian, Yonghong
author_facet Lv, Liuzhenghao
Li, Hao
Wang, Yu
Yan, Zhiyuan
Chen, Zijun
Lin, Zongying
Yuan, Li
Tian, Yonghong
contents Chemical language models (CLMs) are prominent for their effectiveness in exploring chemical space and enabling molecular engineering. However, while exploring chemical-linguistic space, CLMs suffer from the gap between natural language and molecular representations. This challenge is primarily due to the inherent modeling differences between molecules and texts: molecules operate unified modeling to learn chemical space, while natural language sequentially models the semantic space. Additionally, the limited availability of high-quality text-to-molecule datasets further exacerbates this challenge. To address the problem, we first verified the information bias in molecular representations from different perspectives. We then developed the Heterogeneous Molecular Encoding (HME) framework, a unified molecular encoder compressing the molecular features from fragment sequence, topology, and conformation with Q-learning. To better model chemical-linguistic space, we further constructed the MCMoD dataset, which contains over one million molecules with various conditions, including properties, fragments, and descriptions. Experimentally, HME promotes CLMs to achieve chemical-linguistic sharing space exploration: (1) chemical space exploration with linguistic guidance, where HME achieves significant improvements (+8.9\% FCD) for molecular design in multiple constraints, even in zero-shot scenarios; (2) linguistic space exploration with molecular guidance, where HME generates textual descriptions with high qualities (+11.6\% BLEU) for molecules. These results highlight the precision of HME in handling multi-objective and cross-domain tasks, as well as its remarkable generalization capability on unseen task combinations. HME offers a new perspective on navigating chemical-linguistic sharing space, advancing the potential of CLMs in both fundamental research and practical applications in chemistry.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20888
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Navigating Chemical-Linguistic Sharing Space with Heterogeneous Molecular Encoding
Lv, Liuzhenghao
Li, Hao
Wang, Yu
Yan, Zhiyuan
Chen, Zijun
Lin, Zongying
Yuan, Li
Tian, Yonghong
Computational Engineering, Finance, and Science
Chemical language models (CLMs) are prominent for their effectiveness in exploring chemical space and enabling molecular engineering. However, while exploring chemical-linguistic space, CLMs suffer from the gap between natural language and molecular representations. This challenge is primarily due to the inherent modeling differences between molecules and texts: molecules operate unified modeling to learn chemical space, while natural language sequentially models the semantic space. Additionally, the limited availability of high-quality text-to-molecule datasets further exacerbates this challenge. To address the problem, we first verified the information bias in molecular representations from different perspectives. We then developed the Heterogeneous Molecular Encoding (HME) framework, a unified molecular encoder compressing the molecular features from fragment sequence, topology, and conformation with Q-learning. To better model chemical-linguistic space, we further constructed the MCMoD dataset, which contains over one million molecules with various conditions, including properties, fragments, and descriptions. Experimentally, HME promotes CLMs to achieve chemical-linguistic sharing space exploration: (1) chemical space exploration with linguistic guidance, where HME achieves significant improvements (+8.9\% FCD) for molecular design in multiple constraints, even in zero-shot scenarios; (2) linguistic space exploration with molecular guidance, where HME generates textual descriptions with high qualities (+11.6\% BLEU) for molecules. These results highlight the precision of HME in handling multi-objective and cross-domain tasks, as well as its remarkable generalization capability on unseen task combinations. HME offers a new perspective on navigating chemical-linguistic sharing space, advancing the potential of CLMs in both fundamental research and practical applications in chemistry.
title Navigating Chemical-Linguistic Sharing Space with Heterogeneous Molecular Encoding
topic Computational Engineering, Finance, and Science
url https://arxiv.org/abs/2412.20888