Effortless Vision-Language Model Specialization in Histopathology without Annotation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Qiu, Jingna, Jain, Nishanth, Ammeling, Jonas, Aubreville, Marc, Breininger, Katharina
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913984337149952
author Qiu, Jingna
Jain, Nishanth
Ammeling, Jonas
Aubreville, Marc
Breininger, Katharina
author_facet Qiu, Jingna
Jain, Nishanth
Ammeling, Jonas
Aubreville, Marc
Breininger, Katharina
contents Recent advances in Vision-Language Models (VLMs) in histopathology, such as CONCH and QuiltNet, have demonstrated impressive zero-shot classification capabilities across various tasks. However, their general-purpose design may lead to suboptimal performance in specific downstream applications. While supervised fine-tuning methods address this issue, they require manually labeled samples for adaptation. This paper investigates annotation-free adaptation of VLMs through continued pretraining on domain- and task-relevant image-caption pairs extracted from existing databases. Our experiments on two VLMs, CONCH and QuiltNet, across three downstream tasks reveal that these pairs substantially enhance both zero-shot and few-shot performance. Notably, with larger training sizes, continued pretraining matches the performance of few-shot methods while eliminating manual labeling. Its effectiveness, task-agnostic design, and annotation-free workflow make it a promising pathway for adapting VLMs to new histopathology tasks. Code is available at https://github.com/DeepMicroscopy/Annotation-free-VLM-specialization.
format Preprint
id arxiv_https___arxiv_org_abs_2508_07835
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Effortless Vision-Language Model Specialization in Histopathology without Annotation
Qiu, Jingna
Jain, Nishanth
Ammeling, Jonas
Aubreville, Marc
Breininger, Katharina
Computer Vision and Pattern Recognition
Recent advances in Vision-Language Models (VLMs) in histopathology, such as CONCH and QuiltNet, have demonstrated impressive zero-shot classification capabilities across various tasks. However, their general-purpose design may lead to suboptimal performance in specific downstream applications. While supervised fine-tuning methods address this issue, they require manually labeled samples for adaptation. This paper investigates annotation-free adaptation of VLMs through continued pretraining on domain- and task-relevant image-caption pairs extracted from existing databases. Our experiments on two VLMs, CONCH and QuiltNet, across three downstream tasks reveal that these pairs substantially enhance both zero-shot and few-shot performance. Notably, with larger training sizes, continued pretraining matches the performance of few-shot methods while eliminating manual labeling. Its effectiveness, task-agnostic design, and annotation-free workflow make it a promising pathway for adapting VLMs to new histopathology tasks. Code is available at https://github.com/DeepMicroscopy/Annotation-free-VLM-specialization.
title Effortless Vision-Language Model Specialization in Histopathology without Annotation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.07835