Large scale paired antibody language models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kenlay, Henry, Dreyer, Frédéric A., Kovaltsuk, Aleksandr, Miketa, Dom, Pires, Douglas, Deane, Charlotte M.
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909151253233664
author Kenlay, Henry
Dreyer, Frédéric A.
Kovaltsuk, Aleksandr
Miketa, Dom
Pires, Douglas
Deane, Charlotte M.
author_facet Kenlay, Henry
Dreyer, Frédéric A.
Kovaltsuk, Aleksandr
Miketa, Dom
Pires, Douglas
Deane, Charlotte M.
contents Antibodies are proteins produced by the immune system that can identify and neutralise a wide variety of antigens with high specificity and affinity, and constitute the most successful class of biotherapeutics. With the advent of next-generation sequencing, billions of antibody sequences have been collected in recent years, though their application in the design of better therapeutics has been constrained by the sheer volume and complexity of the data. To address this challenge, we present IgBert and IgT5, the best performing antibody-specific language models developed to date which can consistently handle both paired and unpaired variable region sequences as input. These models are trained comprehensively using the more than two billion unpaired sequences and two million paired sequences of light and heavy chains present in the Observed Antibody Space dataset. We show that our models outperform existing antibody and protein language models on a diverse range of design and regression tasks relevant to antibody engineering. This advancement marks a significant leap forward in leveraging machine learning, large scale data sets and high-performance computing for enhancing antibody design for therapeutic development.
format Preprint
id arxiv_https___arxiv_org_abs_2403_17889
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Large scale paired antibody language models
Kenlay, Henry
Dreyer, Frédéric A.
Kovaltsuk, Aleksandr
Miketa, Dom
Pires, Douglas
Deane, Charlotte M.
Biomolecules
Machine Learning
Antibodies are proteins produced by the immune system that can identify and neutralise a wide variety of antigens with high specificity and affinity, and constitute the most successful class of biotherapeutics. With the advent of next-generation sequencing, billions of antibody sequences have been collected in recent years, though their application in the design of better therapeutics has been constrained by the sheer volume and complexity of the data. To address this challenge, we present IgBert and IgT5, the best performing antibody-specific language models developed to date which can consistently handle both paired and unpaired variable region sequences as input. These models are trained comprehensively using the more than two billion unpaired sequences and two million paired sequences of light and heavy chains present in the Observed Antibody Space dataset. We show that our models outperform existing antibody and protein language models on a diverse range of design and regression tasks relevant to antibody engineering. This advancement marks a significant leap forward in leveraging machine learning, large scale data sets and high-performance computing for enhancing antibody design for therapeutic development.
title Large scale paired antibody language models
topic Biomolecules
Machine Learning
url https://arxiv.org/abs/2403.17889