TAPE: A two-stage parameter-efficient adaptation framework for foundation models in OCT-OCTA analysis

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Su, Xiaofei, Wang, Zengshuo, Sun, Minghe, Zhao, Xin, Sun, Mingzhu
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911569357570048
author Su, Xiaofei
Wang, Zengshuo
Sun, Minghe
Zhao, Xin
Sun, Mingzhu
author_facet Su, Xiaofei
Wang, Zengshuo
Sun, Minghe
Zhao, Xin
Sun, Mingzhu
contents Automated analysis of optical coherence tomography (OCT) and OCT angiography (OCTA) images is critical for robust ophthalmic diagnosis. Existing mainstream methods trained from scratch rely heavily on massive data and model scale, thereby hindering their practical deployment in resource-constrained clinical settings. Although transfer learning based on foundation models (FMs) is promising, it still faces significant challenges: domain shift and task misalignment. To address these, we propose TAPE: A Two-stage Adaptation Framework via Parameter-Efficient Fine-tuning, which strategically decouples adaptation into domain alignment and task fitting for downstream segmentation. The domain adaptation stage notably applies parameter-efficient fine-tuning (PEFT) in the context of masked image modeling for medical image domain adaptation, a novel approach to the best of our knowledge. Applying TAPE to retinal layer segmentation on both universal (masked auto-encoder, MAE) and specialized (RETFound) FMs, it demonstrates superior parameter efficiency and achieves state-of-the-art generalization performance across diverse pathologies.
format Preprint
id arxiv_https___arxiv_org_abs_2604_04571
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TAPE: A two-stage parameter-efficient adaptation framework for foundation models in OCT-OCTA analysis
Su, Xiaofei
Wang, Zengshuo
Sun, Minghe
Zhao, Xin
Sun, Mingzhu
Computer Vision and Pattern Recognition
Automated analysis of optical coherence tomography (OCT) and OCT angiography (OCTA) images is critical for robust ophthalmic diagnosis. Existing mainstream methods trained from scratch rely heavily on massive data and model scale, thereby hindering their practical deployment in resource-constrained clinical settings. Although transfer learning based on foundation models (FMs) is promising, it still faces significant challenges: domain shift and task misalignment. To address these, we propose TAPE: A Two-stage Adaptation Framework via Parameter-Efficient Fine-tuning, which strategically decouples adaptation into domain alignment and task fitting for downstream segmentation. The domain adaptation stage notably applies parameter-efficient fine-tuning (PEFT) in the context of masked image modeling for medical image domain adaptation, a novel approach to the best of our knowledge. Applying TAPE to retinal layer segmentation on both universal (masked auto-encoder, MAE) and specialized (RETFound) FMs, it demonstrates superior parameter efficiency and achieves state-of-the-art generalization performance across diverse pathologies.
title TAPE: A two-stage parameter-efficient adaptation framework for foundation models in OCT-OCTA analysis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.04571