PepBenchmark: A Standardized Benchmark for Peptide Machine Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jiahui, Wang, Rouyi, Zhou, Kuangqi, Xiao, Tianshu, Zhu, Lingyan, Min, Yaosen, Wang, Yang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914466714615808
author Zhang, Jiahui
Wang, Rouyi
Zhou, Kuangqi
Xiao, Tianshu
Zhu, Lingyan
Min, Yaosen
Wang, Yang
author_facet Zhang, Jiahui
Wang, Rouyi
Zhou, Kuangqi
Xiao, Tianshu
Zhu, Lingyan
Min, Yaosen
Wang, Yang
contents Peptide therapeutics are widely regarded as the "third generation" of drugs, yet progress in peptide Machine Learning (ML) are hindered by the absence of standardized benchmarks. Here we present PepBenchmark, which unifies datasets, preprocessing, and evaluation protocols for peptide drug discovery. PepBenchmark comprises three components: (1) PepBenchData, a well-curated collection comprising 29 canonical-peptide and 6 non-canonical-peptide datasets across 7 groups, systematically covering key aspects of peptide drug development, representing, to the best of our knowledge, the most comprehensive AI-ready dataset resource to date; (2) PepBenchPipeline, a standardized preprocessing pipeline that ensures consistent dataset cleaning, construction, splitting, and feature transformation, mitigating quality issues common in ad hoc pipelines; and (3) PepBenchLeaderboard, a unified evaluation protocol and leaderboard with strong baselines across 4 major methodological families: Fingerprint-based, GNN-based, PLM-based, and SMILES-based models. Together, PepBenchmark provides the first standardized and comparable foundation for peptide drug discovery, facilitating methodological advances and translation into real-world applications. The data and code are publicly available at https://github.com/ZGCI-AI4S-Pep/PepBenchmark/.
format Preprint
id arxiv_https___arxiv_org_abs_2604_10531
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PepBenchmark: A Standardized Benchmark for Peptide Machine Learning
Zhang, Jiahui
Wang, Rouyi
Zhou, Kuangqi
Xiao, Tianshu
Zhu, Lingyan
Min, Yaosen
Wang, Yang
Machine Learning
Artificial Intelligence
Peptide therapeutics are widely regarded as the "third generation" of drugs, yet progress in peptide Machine Learning (ML) are hindered by the absence of standardized benchmarks. Here we present PepBenchmark, which unifies datasets, preprocessing, and evaluation protocols for peptide drug discovery. PepBenchmark comprises three components: (1) PepBenchData, a well-curated collection comprising 29 canonical-peptide and 6 non-canonical-peptide datasets across 7 groups, systematically covering key aspects of peptide drug development, representing, to the best of our knowledge, the most comprehensive AI-ready dataset resource to date; (2) PepBenchPipeline, a standardized preprocessing pipeline that ensures consistent dataset cleaning, construction, splitting, and feature transformation, mitigating quality issues common in ad hoc pipelines; and (3) PepBenchLeaderboard, a unified evaluation protocol and leaderboard with strong baselines across 4 major methodological families: Fingerprint-based, GNN-based, PLM-based, and SMILES-based models. Together, PepBenchmark provides the first standardized and comparable foundation for peptide drug discovery, facilitating methodological advances and translation into real-world applications. The data and code are publicly available at https://github.com/ZGCI-AI4S-Pep/PepBenchmark/.
title PepBenchmark: A Standardized Benchmark for Peptide Machine Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2604.10531