When Molecular Similarity Works: Property Cliffs Reveal Hidden Errors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Di, Li, Kun, Rao, Haojie, Hu, Longtao, Chen, Jiameng, Hu, Wenbin, Zheng, Yizhen, Yu, Jiajun, Cao, Duanhua
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911691776720896
author Hu, Di
Li, Kun
Rao, Haojie
Hu, Longtao
Chen, Jiameng
Hu, Wenbin
Zheng, Yizhen
Yu, Jiajun
Cao, Duanhua
author_facet Hu, Di
Li, Kun
Rao, Haojie
Hu, Longtao
Chen, Jiameng
Hu, Wenbin
Zheng, Yizhen
Yu, Jiajun
Cao, Duanhua
contents Accurate prediction of molecular properties underpins drug discovery and material design, yet even state-of-the-art models remain vulnerable to localized failure modes that aggregate metrics cannot detect. The places where molecular similarity should be most helpful are also places where standard evaluation can be most misleading. Property cliffs expose this gap: structurally similar molecules can still differ sharply in target property, so models with competitive overall performance may fail in high-risk local neighborhoods. To expose and mitigate this failure mode, CliffSplit, a cliff-aware evaluation protocol that constructs locally supported, cliff-exposed test cases, and CliffLoss, a model-agnostic train-only mitigation mechanism for cliff-sensitive errors, are introduced. Experiments on three QM9 targets and three MoleculeNet tasks across five backbones show that CliffSplit reveals at least 15% higher error in cliff-heavy QM9 regions, while CliffLoss reduces the cliff-to-smooth error gap by up to 30% on Lipophilicity and improves overall MAE by 9.7%. Together, these results turn molecular similarity failure from a descriptive anomaly into a benchmarked evaluation problem for molecular machine learning. The code is available at https://anonymous.4open.science/r/Cliff_Loss.
format Preprint
id arxiv_https___arxiv_org_abs_2605_17265
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When Molecular Similarity Works: Property Cliffs Reveal Hidden Errors
Hu, Di
Li, Kun
Rao, Haojie
Hu, Longtao
Chen, Jiameng
Hu, Wenbin
Zheng, Yizhen
Yu, Jiajun
Cao, Duanhua
Machine Learning
Accurate prediction of molecular properties underpins drug discovery and material design, yet even state-of-the-art models remain vulnerable to localized failure modes that aggregate metrics cannot detect. The places where molecular similarity should be most helpful are also places where standard evaluation can be most misleading. Property cliffs expose this gap: structurally similar molecules can still differ sharply in target property, so models with competitive overall performance may fail in high-risk local neighborhoods. To expose and mitigate this failure mode, CliffSplit, a cliff-aware evaluation protocol that constructs locally supported, cliff-exposed test cases, and CliffLoss, a model-agnostic train-only mitigation mechanism for cliff-sensitive errors, are introduced. Experiments on three QM9 targets and three MoleculeNet tasks across five backbones show that CliffSplit reveals at least 15% higher error in cliff-heavy QM9 regions, while CliffLoss reduces the cliff-to-smooth error gap by up to 30% on Lipophilicity and improves overall MAE by 9.7%. Together, these results turn molecular similarity failure from a descriptive anomaly into a benchmarked evaluation problem for molecular machine learning. The code is available at https://anonymous.4open.science/r/Cliff_Loss.
title When Molecular Similarity Works: Property Cliffs Reveal Hidden Errors
topic Machine Learning
url https://arxiv.org/abs/2605.17265