Driving Accurate Allergen Prediction with Protein Language Models and Generalization-Focused Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wong, Brian Shing-Hei, Kim, Joshua Mincheol, Fung, Sin-Hang, Xiong, Qing, Ao, Kelvin Fu-Kiu, Wei, Junkang, Wang, Ran, Wang, Dan Michelle, Zhou, Jingying, Feng, Bo, Cheng, Alfred Sze-Lok, Yip, Kevin Y., Tsui, Stephen Kwok-Wing, Cao, Qin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912539024031744
author Wong, Brian Shing-Hei
Kim, Joshua Mincheol
Fung, Sin-Hang
Xiong, Qing
Ao, Kelvin Fu-Kiu
Wei, Junkang
Wang, Ran
Wang, Dan Michelle
Zhou, Jingying
Feng, Bo
Cheng, Alfred Sze-Lok
Yip, Kevin Y.
Tsui, Stephen Kwok-Wing
Cao, Qin
author_facet Wong, Brian Shing-Hei
Kim, Joshua Mincheol
Fung, Sin-Hang
Xiong, Qing
Ao, Kelvin Fu-Kiu
Wei, Junkang
Wang, Ran
Wang, Dan Michelle
Zhou, Jingying
Feng, Bo
Cheng, Alfred Sze-Lok
Yip, Kevin Y.
Tsui, Stephen Kwok-Wing
Cao, Qin
contents Allergens, typically proteins capable of triggering adverse immune responses, represent a significant public health challenge. To accurately identify allergen proteins, we introduce Applm (Allergen Prediction with Protein Language Models), a computational framework that leverages the 100-billion parameter xTrimoPGLM protein language model. We show that Applm consistently outperforms seven state-of-the-art methods in a diverse set of tasks that closely resemble difficult real-world scenarios. These include identifying novel allergens that lack similar examples in the training set, differentiating between allergens and non-allergens among homologs with high sequence similarity, and assessing functional consequences of mutations that create few changes to the protein sequences. Our analysis confirms that xTrimoPGLM, originally trained on one trillion tokens to capture general protein sequence characteristics, is crucial for Applm's performance by detecting important differences among protein sequences. In addition to providing Applm as open-source software, we also provide our carefully curated benchmark datasets to facilitate future research.
format Preprint
id arxiv_https___arxiv_org_abs_2508_10541
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Driving Accurate Allergen Prediction with Protein Language Models and Generalization-Focused Evaluation
Wong, Brian Shing-Hei
Kim, Joshua Mincheol
Fung, Sin-Hang
Xiong, Qing
Ao, Kelvin Fu-Kiu
Wei, Junkang
Wang, Ran
Wang, Dan Michelle
Zhou, Jingying
Feng, Bo
Cheng, Alfred Sze-Lok
Yip, Kevin Y.
Tsui, Stephen Kwok-Wing
Cao, Qin
Machine Learning
Quantitative Methods
Allergens, typically proteins capable of triggering adverse immune responses, represent a significant public health challenge. To accurately identify allergen proteins, we introduce Applm (Allergen Prediction with Protein Language Models), a computational framework that leverages the 100-billion parameter xTrimoPGLM protein language model. We show that Applm consistently outperforms seven state-of-the-art methods in a diverse set of tasks that closely resemble difficult real-world scenarios. These include identifying novel allergens that lack similar examples in the training set, differentiating between allergens and non-allergens among homologs with high sequence similarity, and assessing functional consequences of mutations that create few changes to the protein sequences. Our analysis confirms that xTrimoPGLM, originally trained on one trillion tokens to capture general protein sequence characteristics, is crucial for Applm's performance by detecting important differences among protein sequences. In addition to providing Applm as open-source software, we also provide our carefully curated benchmark datasets to facilitate future research.
title Driving Accurate Allergen Prediction with Protein Language Models and Generalization-Focused Evaluation
topic Machine Learning
Quantitative Methods
url https://arxiv.org/abs/2508.10541