Improving Antibody Humanness Prediction using Patent Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ucar, Talip, Ramon, Aubin, Oglic, Dino, Croasdale-Wood, Rebecca, Diethe, Tom, Sormanni, Pietro
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916279872389120
author Ucar, Talip
Ramon, Aubin
Oglic, Dino
Croasdale-Wood, Rebecca
Diethe, Tom
Sormanni, Pietro
author_facet Ucar, Talip
Ramon, Aubin
Oglic, Dino
Croasdale-Wood, Rebecca
Diethe, Tom
Sormanni, Pietro
contents We investigate the potential of patent data for improving the antibody humanness prediction using a multi-stage, multi-loss training process. Humanness serves as a proxy for the immunogenic response to antibody therapeutics, one of the major causes of attrition in drug discovery and a challenging obstacle for their use in clinical settings. We pose the initial learning stage as a weakly-supervised contrastive-learning problem, where each antibody sequence is associated with possibly multiple identifiers of function and the objective is to learn an encoder that groups them according to their patented properties. We then freeze a part of the contrastive encoder and continue training it on the patent data using the cross-entropy loss to predict the humanness score of a given antibody sequence. We illustrate the utility of the patent data and our approach by performing inference on three different immunogenicity datasets, unseen during training. Our empirical results demonstrate that the learned model consistently outperforms the alternative baselines and establishes new state-of-the-art on five out of six inference tasks, irrespective of the used metric.
format Preprint
id arxiv_https___arxiv_org_abs_2401_14442
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Antibody Humanness Prediction using Patent Data
Ucar, Talip
Ramon, Aubin
Oglic, Dino
Croasdale-Wood, Rebecca
Diethe, Tom
Sormanni, Pietro
Quantitative Methods
Machine Learning
We investigate the potential of patent data for improving the antibody humanness prediction using a multi-stage, multi-loss training process. Humanness serves as a proxy for the immunogenic response to antibody therapeutics, one of the major causes of attrition in drug discovery and a challenging obstacle for their use in clinical settings. We pose the initial learning stage as a weakly-supervised contrastive-learning problem, where each antibody sequence is associated with possibly multiple identifiers of function and the objective is to learn an encoder that groups them according to their patented properties. We then freeze a part of the contrastive encoder and continue training it on the patent data using the cross-entropy loss to predict the humanness score of a given antibody sequence. We illustrate the utility of the patent data and our approach by performing inference on three different immunogenicity datasets, unseen during training. Our empirical results demonstrate that the learned model consistently outperforms the alternative baselines and establishes new state-of-the-art on five out of six inference tasks, irrespective of the used metric.
title Improving Antibody Humanness Prediction using Patent Data
topic Quantitative Methods
Machine Learning
url https://arxiv.org/abs/2401.14442