MoleCode unlocks structural intelligence in large language models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yan, Zhiyuan, Liu, Chen, Zhao, Boxuan, Lin, Kaiqing, Zhao, Jixiang, Wang, Yimi, Lv, Liuzhenghao, Li, Hao, Zhang, Shanzhuo, Yuan, Li, Mo, Fanyang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914571213602816
author Yan, Zhiyuan
Liu, Chen
Zhao, Boxuan
Lin, Kaiqing
Zhao, Jixiang
Wang, Yimi
Lv, Liuzhenghao
Li, Hao
Zhang, Shanzhuo
Yuan, Li
Mo, Fanyang
author_facet Yan, Zhiyuan
Liu, Chen
Zhao, Boxuan
Lin, Kaiqing
Zhao, Jixiang
Wang, Yimi
Lv, Liuzhenghao
Li, Hao
Zhang, Shanzhuo
Yuan, Li
Mo, Fanyang
contents Molecules are graphs, but large language models~(LLMs) are usually asked to reason about them through linear strings. The most popular molecular representation, SMILES, compresses atoms, bonds, branches and rings into a compact sequence in which topology is implicit, forcing LLMs to reconstruct molecular structure before performing the requested chemical operation. Here we introduce MoleCode, an LLM-native, training-free, graph-explicit molecular language in which all molecular components are represented as typed entities with persistent identifiers and explicit relations. MoleCode makes molecular topology directly readable, editable and auditable within the language context, allowing an LLM to operate on structure rather than recover it from syntax. Across molecular reasoning, editing, generation and analysis tasks, this representational shift improves frontier LLMs most strongly when structural access is limiting: unfamiliar molecules, topology-sensitive operations, larger structures and repetitive polymers. It also changes how inference is allocated, replacing long reasoning traces devoted to implicit structural reconstruction with shorter, more chemically directed reasoning over explicit atoms and bonds. In molecular optimization, this enables localized, property-aligned edits that preserve structural similarity to the starting compounds. The same Subgraph--Node--Edge grammar extends beyond small molecules to polymers, Markush structures, mechanism-style transformations and interleaved scientific documents, including research articles and patent disclosures in which chemical information is distributed across text and images. These results suggest that the interface between scientific objects and LLMs should not treat structure as something to be decoded from text. When the object of reasoning is relational, the structure itself should be part of the language.
format Preprint
id arxiv_https___arxiv_org_abs_2605_16480
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MoleCode unlocks structural intelligence in large language models
Yan, Zhiyuan
Liu, Chen
Zhao, Boxuan
Lin, Kaiqing
Zhao, Jixiang
Wang, Yimi
Lv, Liuzhenghao
Li, Hao
Zhang, Shanzhuo
Yuan, Li
Mo, Fanyang
Biomolecules
Artificial Intelligence
Molecules are graphs, but large language models~(LLMs) are usually asked to reason about them through linear strings. The most popular molecular representation, SMILES, compresses atoms, bonds, branches and rings into a compact sequence in which topology is implicit, forcing LLMs to reconstruct molecular structure before performing the requested chemical operation. Here we introduce MoleCode, an LLM-native, training-free, graph-explicit molecular language in which all molecular components are represented as typed entities with persistent identifiers and explicit relations. MoleCode makes molecular topology directly readable, editable and auditable within the language context, allowing an LLM to operate on structure rather than recover it from syntax. Across molecular reasoning, editing, generation and analysis tasks, this representational shift improves frontier LLMs most strongly when structural access is limiting: unfamiliar molecules, topology-sensitive operations, larger structures and repetitive polymers. It also changes how inference is allocated, replacing long reasoning traces devoted to implicit structural reconstruction with shorter, more chemically directed reasoning over explicit atoms and bonds. In molecular optimization, this enables localized, property-aligned edits that preserve structural similarity to the starting compounds. The same Subgraph--Node--Edge grammar extends beyond small molecules to polymers, Markush structures, mechanism-style transformations and interleaved scientific documents, including research articles and patent disclosures in which chemical information is distributed across text and images. These results suggest that the interface between scientific objects and LLMs should not treat structure as something to be decoded from text. When the object of reasoning is relational, the structure itself should be part of the language.
title MoleCode unlocks structural intelligence in large language models
topic Biomolecules
Artificial Intelligence
url https://arxiv.org/abs/2605.16480