From Human Labels to Literature: Semi-Supervised Learning of NMR Chemical Shifts at Scale

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Yongqi, Wang, Yecheng, Wang, Jun-jie, Zhu, Rong, Ke, Guolin, E, Weinan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908788891582464
author Jin, Yongqi
Wang, Yecheng
Wang, Jun-jie
Zhu, Rong
Ke, Guolin
E, Weinan
author_facet Jin, Yongqi
Wang, Yecheng
Wang, Jun-jie
Zhu, Rong
Ke, Guolin
E, Weinan
contents Accurate prediction of nuclear magnetic resonance (NMR) chemical shifts is fundamental to spectral analysis and molecular structure elucidation, yet existing machine learning methods rely on limited, labor-intensive atom-assigned datasets. We propose a semi-supervised framework that learns NMR chemical shifts from millions of literature-extracted spectra without explicit atom-level assignments, integrating a small amount of labeled data with large-scale unassigned spectra. We formulate chemical shift prediction from literature spectra as a permutation-invariant set supervision problem, and show that under commonly satisfied conditions on the loss function, optimal bipartite matching reduces to a sorting-based loss, enabling stable large-scale semi-supervised training beyond traditional curated datasets. Our models achieve substantially improved accuracy and robustness over state-of-the-art methods and exhibit stronger generalization on significantly larger and more diverse molecular datasets. Moreover, by incorporating solvent information at scale, our approach captures systematic solvent effects across common NMR solvents for the first time. Overall, our results demonstrate that large-scale unlabeled spectra mined from the literature can serve as a practical and effective data source for training NMR shift models, suggesting a broader role of literature-derived, weakly structured data in data-centric AI for science.
format Preprint
id arxiv_https___arxiv_org_abs_2601_18524
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Human Labels to Literature: Semi-Supervised Learning of NMR Chemical Shifts at Scale
Jin, Yongqi
Wang, Yecheng
Wang, Jun-jie
Zhu, Rong
Ke, Guolin
E, Weinan
Machine Learning
Accurate prediction of nuclear magnetic resonance (NMR) chemical shifts is fundamental to spectral analysis and molecular structure elucidation, yet existing machine learning methods rely on limited, labor-intensive atom-assigned datasets. We propose a semi-supervised framework that learns NMR chemical shifts from millions of literature-extracted spectra without explicit atom-level assignments, integrating a small amount of labeled data with large-scale unassigned spectra. We formulate chemical shift prediction from literature spectra as a permutation-invariant set supervision problem, and show that under commonly satisfied conditions on the loss function, optimal bipartite matching reduces to a sorting-based loss, enabling stable large-scale semi-supervised training beyond traditional curated datasets. Our models achieve substantially improved accuracy and robustness over state-of-the-art methods and exhibit stronger generalization on significantly larger and more diverse molecular datasets. Moreover, by incorporating solvent information at scale, our approach captures systematic solvent effects across common NMR solvents for the first time. Overall, our results demonstrate that large-scale unlabeled spectra mined from the literature can serve as a practical and effective data source for training NMR shift models, suggesting a broader role of literature-derived, weakly structured data in data-centric AI for science.
title From Human Labels to Literature: Semi-Supervised Learning of NMR Chemical Shifts at Scale
topic Machine Learning
url https://arxiv.org/abs/2601.18524