ProtoBERT-LoRA: Parameter-Efficient Prototypical Finetuning for Immunotherapy Study Identification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Shijia, Ding, Xiyu, Ding, Kai, Zhang, Jacob, Galinsky, Kevin, Wang, Mengrui, Mayers, Ryan P., Wang, Zheyu, Kharrazi, Hadi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916662803955712
author Zhang, Shijia
Ding, Xiyu
Ding, Kai
Zhang, Jacob
Galinsky, Kevin
Wang, Mengrui
Mayers, Ryan P.
Wang, Zheyu
Kharrazi, Hadi
author_facet Zhang, Shijia
Ding, Xiyu
Ding, Kai
Zhang, Jacob
Galinsky, Kevin
Wang, Mengrui
Mayers, Ryan P.
Wang, Zheyu
Kharrazi, Hadi
contents Identifying immune checkpoint inhibitor (ICI) studies in genomic repositories like Gene Expression Omnibus (GEO) is vital for cancer research yet remains challenging due to semantic ambiguity, extreme class imbalance, and limited labeled data in low-resource settings. We present ProtoBERT-LoRA, a hybrid framework that combines PubMedBERT with prototypical networks and Low-Rank Adaptation (LoRA) for efficient fine-tuning. The model enforces class-separable embeddings via episodic prototype training while preserving biomedical domain knowledge. Our dataset was divided as: Training (20 positive, 20 negative), Prototype Set (10 positive, 10 negative), Validation (20 positive, 200 negative), and Test (71 positive, 765 negative). Evaluated on test dataset, ProtoBERT-LoRA achieved F1-score of 0.624 (precision: 0.481, recall: 0.887), outperforming the rule-based system, machine learning baselines and finetuned PubMedBERT. Application to 44,287 unlabeled studies reduced manual review efforts by 82%. Ablation studies confirmed that combining prototypes with LoRA improved performance by 29% over stand-alone LoRA.
format Preprint
id arxiv_https___arxiv_org_abs_2503_20179
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ProtoBERT-LoRA: Parameter-Efficient Prototypical Finetuning for Immunotherapy Study Identification
Zhang, Shijia
Ding, Xiyu
Ding, Kai
Zhang, Jacob
Galinsky, Kevin
Wang, Mengrui
Mayers, Ryan P.
Wang, Zheyu
Kharrazi, Hadi
Computation and Language
Information Retrieval
Quantitative Methods
Identifying immune checkpoint inhibitor (ICI) studies in genomic repositories like Gene Expression Omnibus (GEO) is vital for cancer research yet remains challenging due to semantic ambiguity, extreme class imbalance, and limited labeled data in low-resource settings. We present ProtoBERT-LoRA, a hybrid framework that combines PubMedBERT with prototypical networks and Low-Rank Adaptation (LoRA) for efficient fine-tuning. The model enforces class-separable embeddings via episodic prototype training while preserving biomedical domain knowledge. Our dataset was divided as: Training (20 positive, 20 negative), Prototype Set (10 positive, 10 negative), Validation (20 positive, 200 negative), and Test (71 positive, 765 negative). Evaluated on test dataset, ProtoBERT-LoRA achieved F1-score of 0.624 (precision: 0.481, recall: 0.887), outperforming the rule-based system, machine learning baselines and finetuned PubMedBERT. Application to 44,287 unlabeled studies reduced manual review efforts by 82%. Ablation studies confirmed that combining prototypes with LoRA improved performance by 29% over stand-alone LoRA.
title ProtoBERT-LoRA: Parameter-Efficient Prototypical Finetuning for Immunotherapy Study Identification
topic Computation and Language
Information Retrieval
Quantitative Methods
url https://arxiv.org/abs/2503.20179