Modular Multi-Task Learning for Chemical Reaction Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pang, Jiayun, Zaitoun, Ahmed M., Cambeiro, Xacobe Couso, Vulić, Ivan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908827494907904
author Pang, Jiayun
Zaitoun, Ahmed M.
Cambeiro, Xacobe Couso
Vulić, Ivan
author_facet Pang, Jiayun
Zaitoun, Ahmed M.
Cambeiro, Xacobe Couso
Vulić, Ivan
contents Adapting large language models (LLMs) trained on broad organic chemistry to smaller, domain-specific reaction datasets is a key challenge in chemical and pharmaceutical R&D. Effective specialisation requires learning new reaction knowledge while preserving general chemical understanding across related tasks. Here, we evaluate Low-Rank Adaptation (LoRA) as a parameter-efficient alternative to full fine-tuning for organic reaction prediction on limited, complex datasets. Using USPTO reaction classes and challenging C-H functionalisation reactions, we benchmark forward reaction prediction, retrosynthesis and reagent prediction. LoRA achieves accuracy comparable to full fine-tuning while effectively mitigating catastrophic forgetting and better preserving multi-task performance. Both fine-tuning approaches generalise beyond training distributions, producing plausible alternative solvent predictions. Notably, C-H functionalisation fine-tuning reveals that LoRA and full fine-tuning encode subtly different reactivity patterns, suggesting more effective reaction-specific adaptation with LoRA. As LLMs continue to scale, our results highlight the practicality of modular, parameter-efficient fine-tuning strategies for their flexible deployment for chemistry applications.
format Preprint
id arxiv_https___arxiv_org_abs_2602_10404
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Modular Multi-Task Learning for Chemical Reaction Prediction
Pang, Jiayun
Zaitoun, Ahmed M.
Cambeiro, Xacobe Couso
Vulić, Ivan
Machine Learning
Artificial Intelligence
Computation and Language
Adapting large language models (LLMs) trained on broad organic chemistry to smaller, domain-specific reaction datasets is a key challenge in chemical and pharmaceutical R&D. Effective specialisation requires learning new reaction knowledge while preserving general chemical understanding across related tasks. Here, we evaluate Low-Rank Adaptation (LoRA) as a parameter-efficient alternative to full fine-tuning for organic reaction prediction on limited, complex datasets. Using USPTO reaction classes and challenging C-H functionalisation reactions, we benchmark forward reaction prediction, retrosynthesis and reagent prediction. LoRA achieves accuracy comparable to full fine-tuning while effectively mitigating catastrophic forgetting and better preserving multi-task performance. Both fine-tuning approaches generalise beyond training distributions, producing plausible alternative solvent predictions. Notably, C-H functionalisation fine-tuning reveals that LoRA and full fine-tuning encode subtly different reactivity patterns, suggesting more effective reaction-specific adaptation with LoRA. As LLMs continue to scale, our results highlight the practicality of modular, parameter-efficient fine-tuning strategies for their flexible deployment for chemistry applications.
title Modular Multi-Task Learning for Chemical Reaction Prediction
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2602.10404