RTMol: Rethinking Molecule-text Alignment in a Round-trip View

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Letian, Shi, Runhan, Yu, Gufeng, Yang, Yang
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918216491597824
author Chen, Letian
Shi, Runhan
Yu, Gufeng
Yang, Yang
author_facet Chen, Letian
Shi, Runhan
Yu, Gufeng
Yang, Yang
contents Aligning molecular sequence representations (e.g., SMILES notations) with textual descriptions is critical for applications spanning drug discovery, materials design, and automated chemical literature analysis. Existing methodologies typically treat molecular captioning (molecule-to-text) and text-based molecular design (text-to-molecule) as separate tasks, relying on supervised fine-tuning or contrastive learning pipelines. These approaches face three key limitations: (i) conventional metrics like BLEU prioritize linguistic fluency over chemical accuracy, (ii) training datasets frequently contain chemically ambiguous narratives with incomplete specifications, and (iii) independent optimization of generation directions leads to bidirectional inconsistency. To address these issues, we propose RTMol, a bidirectional alignment framework that unifies molecular captioning and text-to-SMILES generation through self-supervised round-trip learning. The framework introduces novel round-trip evaluation metrics and enables unsupervised training for molecular captioning without requiring paired molecule-text corpora. Experiments demonstrate that RTMol enhances bidirectional alignment performance by up to 47% across various LLMs, establishing an effective paradigm for joint molecule-text understanding and generation.
format Preprint
id arxiv_https___arxiv_org_abs_2511_12135
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RTMol: Rethinking Molecule-text Alignment in a Round-trip View
Chen, Letian
Shi, Runhan
Yu, Gufeng
Yang, Yang
Artificial Intelligence
Machine Learning
Biomolecules
Aligning molecular sequence representations (e.g., SMILES notations) with textual descriptions is critical for applications spanning drug discovery, materials design, and automated chemical literature analysis. Existing methodologies typically treat molecular captioning (molecule-to-text) and text-based molecular design (text-to-molecule) as separate tasks, relying on supervised fine-tuning or contrastive learning pipelines. These approaches face three key limitations: (i) conventional metrics like BLEU prioritize linguistic fluency over chemical accuracy, (ii) training datasets frequently contain chemically ambiguous narratives with incomplete specifications, and (iii) independent optimization of generation directions leads to bidirectional inconsistency. To address these issues, we propose RTMol, a bidirectional alignment framework that unifies molecular captioning and text-to-SMILES generation through self-supervised round-trip learning. The framework introduces novel round-trip evaluation metrics and enables unsupervised training for molecular captioning without requiring paired molecule-text corpora. Experiments demonstrate that RTMol enhances bidirectional alignment performance by up to 47% across various LLMs, establishing an effective paradigm for joint molecule-text understanding and generation.
title RTMol: Rethinking Molecule-text Alignment in a Round-trip View
topic Artificial Intelligence
Machine Learning
Biomolecules
url https://arxiv.org/abs/2511.12135