Saved in:
Bibliographic Details
Main Authors: Zhang, Yu, Han, Yang, Chen, Shuai, Yu, Ruijie, Zhao, Xin, Liu, Xianbin, Zeng, Kaipeng, Yu, Mengdi, Tian, Jidong, Zhu, Feng, Yang, Xiaokang, Jin, Yaohui, Xu, Yanyan
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2504.18340
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916707592830976
author Zhang, Yu
Han, Yang
Chen, Shuai
Yu, Ruijie
Zhao, Xin
Liu, Xianbin
Zeng, Kaipeng
Yu, Mengdi
Tian, Jidong
Zhu, Feng
Yang, Xiaokang
Jin, Yaohui
Xu, Yanyan
author_facet Zhang, Yu
Han, Yang
Chen, Shuai
Yu, Ruijie
Zhao, Xin
Liu, Xianbin
Zeng, Kaipeng
Yu, Mengdi
Tian, Jidong
Zhu, Feng
Yang, Xiaokang
Jin, Yaohui
Xu, Yanyan
contents Chemical synthesis, as a foundational methodology in the creation of transformative molecules, exerts substantial influence across diverse sectors from life sciences to materials and energy. Current chemical synthesis practices emphasize laborious and costly trial-and-error workflows, underscoring the urgent need for advanced AI assistants. Nowadays, large language models (LLMs), typified by GPT-4, have been introduced as an efficient tool to facilitate scientific research. Here, we present Chemma, a fully fine-tuned LLM with 1.28 million pairs of Q&A about reactions, as an assistant to accelerate organic chemistry synthesis. Chemma surpasses the best-known results in multiple chemical tasks, e.g., single-step retrosynthesis and yield prediction, which highlights the potential of general AI for organic chemistry. Via predicting yields across the experimental reaction space, Chemma significantly improves the reaction exploration capability of Bayesian optimization. More importantly, integrated in an active learning framework, Chemma exhibits advanced potential for autonomous experimental exploration and optimization in open reaction spaces. For an unreported Suzuki-Miyaura cross-coupling reaction of cyclic aminoboronates and aryl halides for the synthesis of $α$-Aryl N-heterocycles, the human-AI collaboration successfully explored suitable ligand and solvent (1,4-dioxane) within only 15 runs, achieving an isolated yield of 67%. These results reveal that, without quantum-chemical calculations, Chemma can comprehend and extract chemical insights from reaction data, in a manner akin to human experts. This work opens avenues for accelerating organic chemistry synthesis with adapted large language models.
format Preprint
id arxiv_https___arxiv_org_abs_2504_18340
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Large Language Models to Accelerate Organic Chemistry Synthesis
Zhang, Yu
Han, Yang
Chen, Shuai
Yu, Ruijie
Zhao, Xin
Liu, Xianbin
Zeng, Kaipeng
Yu, Mengdi
Tian, Jidong
Zhu, Feng
Yang, Xiaokang
Jin, Yaohui
Xu, Yanyan
Chemical Physics
Chemical synthesis, as a foundational methodology in the creation of transformative molecules, exerts substantial influence across diverse sectors from life sciences to materials and energy. Current chemical synthesis practices emphasize laborious and costly trial-and-error workflows, underscoring the urgent need for advanced AI assistants. Nowadays, large language models (LLMs), typified by GPT-4, have been introduced as an efficient tool to facilitate scientific research. Here, we present Chemma, a fully fine-tuned LLM with 1.28 million pairs of Q&A about reactions, as an assistant to accelerate organic chemistry synthesis. Chemma surpasses the best-known results in multiple chemical tasks, e.g., single-step retrosynthesis and yield prediction, which highlights the potential of general AI for organic chemistry. Via predicting yields across the experimental reaction space, Chemma significantly improves the reaction exploration capability of Bayesian optimization. More importantly, integrated in an active learning framework, Chemma exhibits advanced potential for autonomous experimental exploration and optimization in open reaction spaces. For an unreported Suzuki-Miyaura cross-coupling reaction of cyclic aminoboronates and aryl halides for the synthesis of $α$-Aryl N-heterocycles, the human-AI collaboration successfully explored suitable ligand and solvent (1,4-dioxane) within only 15 runs, achieving an isolated yield of 67%. These results reveal that, without quantum-chemical calculations, Chemma can comprehend and extract chemical insights from reaction data, in a manner akin to human experts. This work opens avenues for accelerating organic chemistry synthesis with adapted large language models.
title Large Language Models to Accelerate Organic Chemistry Synthesis
topic Chemical Physics
url https://arxiv.org/abs/2504.18340