L+M-24: Building a Dataset for Language + Molecules @ ACL 2024

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Edwards, Carl, Wang, Qingyun, Zhao, Lawrence, Ji, Heng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909242136461312
author Edwards, Carl
Wang, Qingyun
Zhao, Lawrence
Ji, Heng
author_facet Edwards, Carl
Wang, Qingyun
Zhao, Lawrence
Ji, Heng
contents Language-molecule models have emerged as an exciting direction for molecular discovery and understanding. However, training these models is challenging due to the scarcity of molecule-language pair datasets. At this point, datasets have been released which are 1) small and scraped from existing databases, 2) large but noisy and constructed by performing entity linking on the scientific literature, and 3) built by converting property prediction datasets to natural language using templates. In this document, we detail the $\textit{L+M-24}$ dataset, which has been created for the Language + Molecules Workshop shared task at ACL 2024. In particular, $\textit{L+M-24}$ is designed to focus on three key benefits of natural language in molecule design: compositionality, functionality, and abstraction.
format Preprint
id arxiv_https___arxiv_org_abs_2403_00791
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle L+M-24: Building a Dataset for Language + Molecules @ ACL 2024
Edwards, Carl
Wang, Qingyun
Zhao, Lawrence
Ji, Heng
Computation and Language
Artificial Intelligence
Biomolecules
Quantitative Methods
Language-molecule models have emerged as an exciting direction for molecular discovery and understanding. However, training these models is challenging due to the scarcity of molecule-language pair datasets. At this point, datasets have been released which are 1) small and scraped from existing databases, 2) large but noisy and constructed by performing entity linking on the scientific literature, and 3) built by converting property prediction datasets to natural language using templates. In this document, we detail the $\textit{L+M-24}$ dataset, which has been created for the Language + Molecules Workshop shared task at ACL 2024. In particular, $\textit{L+M-24}$ is designed to focus on three key benefits of natural language in molecule design: compositionality, functionality, and abstraction.
title L+M-24: Building a Dataset for Language + Molecules @ ACL 2024
topic Computation and Language
Artificial Intelligence
Biomolecules
Quantitative Methods
url https://arxiv.org/abs/2403.00791