Continued domain-specific pre-training of protein language models for pMHC-I binding prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mares, Sergio E., Weinberger, Ariel Espinoza, Ioannidis, Nilah M.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916848599040000
author Mares, Sergio E.
Weinberger, Ariel Espinoza
Ioannidis, Nilah M.
author_facet Mares, Sergio E.
Weinberger, Ariel Espinoza
Ioannidis, Nilah M.
contents Predicting peptide--major histocompatibility complex I (pMHC-I) binding affinity remains challenging due to extreme allelic diversity ($\sim$30,000 HLA alleles), severe data scarcity for most alleles, and noisy experimental measurements. Current methods particularly struggle with underrepresented alleles and quantitative binding prediction. We test whether domain-specific continued pre-training of protein language models is beneficial for their application to pMHC-I binding affinity prediction. Starting from ESM Cambrian (300M parameters), we perform masked-language modeling (MLM)-based continued pre-training on HLA-associated peptides (epitopes), testing two input formats: epitope sequences alone versus epitopes concatenated with HLA heavy chain sequences. We then fine-tune for functional IC$_{50}$ binding affinity prediction using only high-quality quantitative data, avoiding mass spectrometry biases that are inherited by existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2507_13077
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Continued domain-specific pre-training of protein language models for pMHC-I binding prediction
Mares, Sergio E.
Weinberger, Ariel Espinoza
Ioannidis, Nilah M.
Quantitative Methods
Predicting peptide--major histocompatibility complex I (pMHC-I) binding affinity remains challenging due to extreme allelic diversity ($\sim$30,000 HLA alleles), severe data scarcity for most alleles, and noisy experimental measurements. Current methods particularly struggle with underrepresented alleles and quantitative binding prediction. We test whether domain-specific continued pre-training of protein language models is beneficial for their application to pMHC-I binding affinity prediction. Starting from ESM Cambrian (300M parameters), we perform masked-language modeling (MLM)-based continued pre-training on HLA-associated peptides (epitopes), testing two input formats: epitope sequences alone versus epitopes concatenated with HLA heavy chain sequences. We then fine-tune for functional IC$_{50}$ binding affinity prediction using only high-quality quantitative data, avoiding mass spectrometry biases that are inherited by existing methods.
title Continued domain-specific pre-training of protein language models for pMHC-I binding prediction
topic Quantitative Methods
url https://arxiv.org/abs/2507.13077