polyRETRO: a Language Model Approach to predict Polymerization Class and Monomer(s) for a Target Polymer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Agarwal, Sakshi, Xiong, Wei, Ramprasad, Rampi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918232253792256
author Agarwal, Sakshi
Xiong, Wei
Ramprasad, Rampi
author_facet Agarwal, Sakshi
Xiong, Wei
Ramprasad, Rampi
contents While machine learning has transformed polymer design by enabling rapid property prediction and candidate generation, translating these designs into experimentally realizable materials remains a critical challenge. Traditionally, the synthesis of target polymers has relied heavily on expert intuition and prior experience. The lack of automated retrosynthetic tools to assist chemists, limit the rapid practical impact of data-driven polymer discovery. To expedite lab-scale validation and beyond, we present a retrosynthetic framework that leverages large language models (LLMs) to guide polymer synthesis. Our approach, which we call polyRETRO, involves two key steps: 1) predicting the most likely polymerization reaction class of a target polymer and 2) identifying the underlying chemical transformation templates and the corresponding monomers, using primarily natural-language based constructs. This LLM-driven framework enables direct retrosynthetic analysis given just the target polymer SMILES string. polyRETRO constitutes a initial step towards a scalable, interpretable, and generalizable approach to bridge the gap between computational design and experimental synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2512_05138
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle polyRETRO: a Language Model Approach to predict Polymerization Class and Monomer(s) for a Target Polymer
Agarwal, Sakshi
Xiong, Wei
Ramprasad, Rampi
Soft Condensed Matter
Materials Science
While machine learning has transformed polymer design by enabling rapid property prediction and candidate generation, translating these designs into experimentally realizable materials remains a critical challenge. Traditionally, the synthesis of target polymers has relied heavily on expert intuition and prior experience. The lack of automated retrosynthetic tools to assist chemists, limit the rapid practical impact of data-driven polymer discovery. To expedite lab-scale validation and beyond, we present a retrosynthetic framework that leverages large language models (LLMs) to guide polymer synthesis. Our approach, which we call polyRETRO, involves two key steps: 1) predicting the most likely polymerization reaction class of a target polymer and 2) identifying the underlying chemical transformation templates and the corresponding monomers, using primarily natural-language based constructs. This LLM-driven framework enables direct retrosynthetic analysis given just the target polymer SMILES string. polyRETRO constitutes a initial step towards a scalable, interpretable, and generalizable approach to bridge the gap between computational design and experimental synthesis.
title polyRETRO: a Language Model Approach to predict Polymerization Class and Monomer(s) for a Target Polymer
topic Soft Condensed Matter
Materials Science
url https://arxiv.org/abs/2512.05138