Saved in:
Bibliographic Details
Main Authors: Hakim, Sheikh Azizul, Roy, Kowshic, Rahman, M Saifur
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.10655
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909840214851584
author Hakim, Sheikh Azizul
Roy, Kowshic
Rahman, M Saifur
author_facet Hakim, Sheikh Azizul
Roy, Kowshic
Rahman, M Saifur
contents Large pretrained language models have transformed natural language processing, and their adaptation to protein sequences -- viewed as strings of amino acid characters -- has advanced protein analysis. However, the distinct properties of proteins, such as variable sequence lengths and lack of word-sentence analogs, necessitate a deeper understanding of protein language models (LMs). We investigate the isotropy of protein LM embedding spaces using average pairwise cosine similarity and the IsoScore method, revealing that models like ProtBERT and ProtXLNet are highly anisotropic, utilizing only 2--14 dimensions for global and local representations. In contrast, multi-modal training in ProteinBERT, which integrates sequence and gene ontology data, enhances isotropy, suggesting that diverse biological inputs improve representational efficiency. We also find that embedding distances weakly correlate with alignment-based similarity scores, particularly at low similarity.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10655
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Isotropy and Geometry of Pretrained Protein LMs
Hakim, Sheikh Azizul
Roy, Kowshic
Rahman, M Saifur
Other Quantitative Biology
Large pretrained language models have transformed natural language processing, and their adaptation to protein sequences -- viewed as strings of amino acid characters -- has advanced protein analysis. However, the distinct properties of proteins, such as variable sequence lengths and lack of word-sentence analogs, necessitate a deeper understanding of protein language models (LMs). We investigate the isotropy of protein LM embedding spaces using average pairwise cosine similarity and the IsoScore method, revealing that models like ProtBERT and ProtXLNet are highly anisotropic, utilizing only 2--14 dimensions for global and local representations. In contrast, multi-modal training in ProteinBERT, which integrates sequence and gene ontology data, enhances isotropy, suggesting that diverse biological inputs improve representational efficiency. We also find that embedding distances weakly correlate with alignment-based similarity scores, particularly at low similarity.
title Isotropy and Geometry of Pretrained Protein LMs
topic Other Quantitative Biology
url https://arxiv.org/abs/2510.10655