Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Shih-Wen, Fan, Hsuan-Yu, Chu, Wei-Ta, Yang, Fu-En, Wang, Yu-Chiang Frank
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912443878342656
author Liu, Shih-Wen
Fan, Hsuan-Yu
Chu, Wei-Ta
Yang, Fu-En
Wang, Yu-Chiang Frank
author_facet Liu, Shih-Wen
Fan, Hsuan-Yu
Chu, Wei-Ta
Yang, Fu-En
Wang, Yu-Chiang Frank
contents Automating medical report generation from histopathology images is a critical challenge requiring effective visual representations and domain-specific knowledge. Inspired by the common practices of human experts, we propose an in-context learning framework called PathGenIC that integrates context derived from the training set with a multimodal in-context learning (ICL) mechanism. Our method dynamically retrieves semantically similar whole slide image (WSI)-report pairs and incorporates adaptive feedback to enhance contextual relevance and generation quality. Evaluated on the HistGen benchmark, the framework achieves state-of-the-art results, with significant improvements across BLEU, METEOR, and ROUGE-L metrics, and demonstrates robustness across diverse report lengths and disease categories. By maximizing training data utility and bridging vision and language with ICL, our work offers a solution for AI-driven histopathology reporting, setting a strong foundation for future advancements in multimodal clinical applications.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17645
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning
Liu, Shih-Wen
Fan, Hsuan-Yu
Chu, Wei-Ta
Yang, Fu-En
Wang, Yu-Chiang Frank
Computer Vision and Pattern Recognition
Automating medical report generation from histopathology images is a critical challenge requiring effective visual representations and domain-specific knowledge. Inspired by the common practices of human experts, we propose an in-context learning framework called PathGenIC that integrates context derived from the training set with a multimodal in-context learning (ICL) mechanism. Our method dynamically retrieves semantically similar whole slide image (WSI)-report pairs and incorporates adaptive feedback to enhance contextual relevance and generation quality. Evaluated on the HistGen benchmark, the framework achieves state-of-the-art results, with significant improvements across BLEU, METEOR, and ROUGE-L metrics, and demonstrates robustness across diverse report lengths and disease categories. By maximizing training data utility and bridging vision and language with ICL, our work offers a solution for AI-driven histopathology reporting, setting a strong foundation for future advancements in multimodal clinical applications.
title Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.17645