Democratising Pathology Co-Pilots: An Open Pipeline and Dataset for Whole-Slide Vision-Language Modelling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Moonemans, Sander, Ram, Sebastiaan, Meeuwsen, Frédérique, Lems, Carlijn, van der Laak, Jeroen, Litjens, Geert, Ciompi, Francesco
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913034263330816
author Moonemans, Sander
Ram, Sebastiaan
Meeuwsen, Frédérique
Lems, Carlijn
van der Laak, Jeroen
Litjens, Geert
Ciompi, Francesco
author_facet Moonemans, Sander
Ram, Sebastiaan
Meeuwsen, Frédérique
Lems, Carlijn
van der Laak, Jeroen
Litjens, Geert
Ciompi, Francesco
contents Vision-language models (VLMs) have the potential to become co-pilots for pathologists. However, most VLMs either focus on small regions of interest within whole-slide images, provide only static slide-level outputs, or rely on data that is not publicly available, limiting reproducibility. Furthermore, training data containing WSIs paired with detailed clinical reports is scarce, restricting progress toward transparent and generalisable VLMs. We address these limitations with three main contributions. First, we introduce Polysome, a standardised tool for synthetic instruction generation. Second, we apply Polysome to the public HISTAI dataset, generating HISTAI-Instruct, a large whole-slide instruction tuning dataset spanning 24,259 slides and over 1.1 million instruction-response pairs. Finally, we use HISTAI-Instruct to train ANTONI-α, a VLM capable of visual-question answering (VQA). We show that ANTONI-α outperforms MedGemma on WSI-level VQA tasks of tissue identification, neoplasm detection, and differential diagnosis. We also compare the performance of multiple incarnations of ANTONI-α trained with different amounts of data. All methods, data, and code are publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17326
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Democratising Pathology Co-Pilots: An Open Pipeline and Dataset for Whole-Slide Vision-Language Modelling
Moonemans, Sander
Ram, Sebastiaan
Meeuwsen, Frédérique
Lems, Carlijn
van der Laak, Jeroen
Litjens, Geert
Ciompi, Francesco
Computer Vision and Pattern Recognition
Vision-language models (VLMs) have the potential to become co-pilots for pathologists. However, most VLMs either focus on small regions of interest within whole-slide images, provide only static slide-level outputs, or rely on data that is not publicly available, limiting reproducibility. Furthermore, training data containing WSIs paired with detailed clinical reports is scarce, restricting progress toward transparent and generalisable VLMs. We address these limitations with three main contributions. First, we introduce Polysome, a standardised tool for synthetic instruction generation. Second, we apply Polysome to the public HISTAI dataset, generating HISTAI-Instruct, a large whole-slide instruction tuning dataset spanning 24,259 slides and over 1.1 million instruction-response pairs. Finally, we use HISTAI-Instruct to train ANTONI-α, a VLM capable of visual-question answering (VQA). We show that ANTONI-α outperforms MedGemma on WSI-level VQA tasks of tissue identification, neoplasm detection, and differential diagnosis. We also compare the performance of multiple incarnations of ANTONI-α trained with different amounts of data. All methods, data, and code are publicly available.
title Democratising Pathology Co-Pilots: An Open Pipeline and Dataset for Whole-Slide Vision-Language Modelling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.17326