Saved in:
Bibliographic Details
Main Authors: Mengren, Liu, Zhang, Yixiang, Yiming, Zhang
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2512.09894
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909955182821376
author Mengren
Liu
Zhang, Yixiang
Yiming
Zhang
author_facet Mengren
Liu
Zhang, Yixiang
Yiming
Zhang
contents Recent advances in protein language models (PLMs) have demonstrated remarkable capabilities in understanding protein sequences. However, the extent to which different model architectures capture antibody-specific biological properties remains unexplored. In this work, we systematically investigate how architectural choices in PLMs influence their ability to comprehend antibody sequence characteristics and functions. We evaluate three state-of-the-art PLMs-AntiBERTa, BioBERT, and ESM2--against a general-purpose language model (GPT-2) baseline on antibody target specificity prediction tasks. Our results demonstrate that while all PLMs achieve high classification accuracy, they exhibit distinct biases in capturing biological features such as V gene usage, somatic hypermutation patterns, and isotype information. Through attention attribution analysis, we show that antibody-specific models like AntiBERTa naturally learn to focus on complementarity-determining regions (CDRs), while general protein models benefit significantly from explicit CDR-focused training strategies. These findings provide insights into the relationship between model architecture and biological feature extraction, offering valuable guidance for future PLM development in computational antibody design.
format Preprint
id arxiv_https___arxiv_org_abs_2512_09894
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring Protein Language Model Architecture-Induced Biases for Antibody Comprehension
Mengren
Liu
Zhang, Yixiang
Yiming
Zhang
Machine Learning
Recent advances in protein language models (PLMs) have demonstrated remarkable capabilities in understanding protein sequences. However, the extent to which different model architectures capture antibody-specific biological properties remains unexplored. In this work, we systematically investigate how architectural choices in PLMs influence their ability to comprehend antibody sequence characteristics and functions. We evaluate three state-of-the-art PLMs-AntiBERTa, BioBERT, and ESM2--against a general-purpose language model (GPT-2) baseline on antibody target specificity prediction tasks. Our results demonstrate that while all PLMs achieve high classification accuracy, they exhibit distinct biases in capturing biological features such as V gene usage, somatic hypermutation patterns, and isotype information. Through attention attribution analysis, we show that antibody-specific models like AntiBERTa naturally learn to focus on complementarity-determining regions (CDRs), while general protein models benefit significantly from explicit CDR-focused training strategies. These findings provide insights into the relationship between model architecture and biological feature extraction, offering valuable guidance for future PLM development in computational antibody design.
title Exploring Protein Language Model Architecture-Induced Biases for Antibody Comprehension
topic Machine Learning
url https://arxiv.org/abs/2512.09894