The Catechol Benchmark: Time-series Solvent Selection Data for Few-shot Machine Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Boyne, Toby, Campos, Juan S., Langdon, Becky D., Qing, Jixiang, Xie, Yilin, Zhang, Shiqiang, Tsay, Calvin, Misener, Ruth, Davies, Daniel W., Jelfs, Kim E., Boyall, Sarah, Dixon, Thomas M., Schrecker, Linden, Folch, Jose Pablo
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912733507616768
author Boyne, Toby
Campos, Juan S.
Langdon, Becky D.
Qing, Jixiang
Xie, Yilin
Zhang, Shiqiang
Tsay, Calvin
Misener, Ruth
Davies, Daniel W.
Jelfs, Kim E.
Boyall, Sarah
Dixon, Thomas M.
Schrecker, Linden
Folch, Jose Pablo
author_facet Boyne, Toby
Campos, Juan S.
Langdon, Becky D.
Qing, Jixiang
Xie, Yilin
Zhang, Shiqiang
Tsay, Calvin
Misener, Ruth
Davies, Daniel W.
Jelfs, Kim E.
Boyall, Sarah
Dixon, Thomas M.
Schrecker, Linden
Folch, Jose Pablo
contents Machine learning has promised to change the landscape of laboratory chemistry, with impressive results in molecular property prediction and reaction retro-synthesis. However, chemical datasets are often inaccessible to the machine learning community as they tend to require cleaning, thorough understanding of the chemistry, or are simply not available. In this paper, we introduce a novel dataset for yield prediction, providing the first-ever transient flow dataset for machine learning benchmarking, covering over 1200 process conditions. While previous datasets focus on discrete parameters, our experimental set-up allow us to sample a large number of continuous process conditions, generating new challenges for machine learning models. We focus on solvent selection, a task that is particularly difficult to model theoretically and therefore ripe for machine learning applications. We showcase benchmarking for regression algorithms, transfer-learning approaches, feature engineering, and active learning, with important applications towards solvent replacement and sustainable manufacturing.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07619
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Catechol Benchmark: Time-series Solvent Selection Data for Few-shot Machine Learning
Boyne, Toby
Campos, Juan S.
Langdon, Becky D.
Qing, Jixiang
Xie, Yilin
Zhang, Shiqiang
Tsay, Calvin
Misener, Ruth
Davies, Daniel W.
Jelfs, Kim E.
Boyall, Sarah
Dixon, Thomas M.
Schrecker, Linden
Folch, Jose Pablo
Machine Learning
Quantitative Methods
Machine learning has promised to change the landscape of laboratory chemistry, with impressive results in molecular property prediction and reaction retro-synthesis. However, chemical datasets are often inaccessible to the machine learning community as they tend to require cleaning, thorough understanding of the chemistry, or are simply not available. In this paper, we introduce a novel dataset for yield prediction, providing the first-ever transient flow dataset for machine learning benchmarking, covering over 1200 process conditions. While previous datasets focus on discrete parameters, our experimental set-up allow us to sample a large number of continuous process conditions, generating new challenges for machine learning models. We focus on solvent selection, a task that is particularly difficult to model theoretically and therefore ripe for machine learning applications. We showcase benchmarking for regression algorithms, transfer-learning approaches, feature engineering, and active learning, with important applications towards solvent replacement and sustainable manufacturing.
title The Catechol Benchmark: Time-series Solvent Selection Data for Few-shot Machine Learning
topic Machine Learning
Quantitative Methods
url https://arxiv.org/abs/2506.07619