Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912443878342656 |
|---|---|
| author | Liu, Shih-Wen Fan, Hsuan-Yu Chu, Wei-Ta Yang, Fu-En Wang, Yu-Chiang Frank |
| author_facet | Liu, Shih-Wen Fan, Hsuan-Yu Chu, Wei-Ta Yang, Fu-En Wang, Yu-Chiang Frank |
| contents | Automating medical report generation from histopathology images is a critical challenge requiring effective visual representations and domain-specific knowledge. Inspired by the common practices of human experts, we propose an in-context learning framework called PathGenIC that integrates context derived from the training set with a multimodal in-context learning (ICL) mechanism. Our method dynamically retrieves semantically similar whole slide image (WSI)-report pairs and incorporates adaptive feedback to enhance contextual relevance and generation quality. Evaluated on the HistGen benchmark, the framework achieves state-of-the-art results, with significant improvements across BLEU, METEOR, and ROUGE-L metrics, and demonstrates robustness across diverse report lengths and disease categories. By maximizing training data utility and bridging vision and language with ICL, our work offers a solution for AI-driven histopathology reporting, setting a strong foundation for future advancements in multimodal clinical applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_17645 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning Liu, Shih-Wen Fan, Hsuan-Yu Chu, Wei-Ta Yang, Fu-En Wang, Yu-Chiang Frank Computer Vision and Pattern Recognition Automating medical report generation from histopathology images is a critical challenge requiring effective visual representations and domain-specific knowledge. Inspired by the common practices of human experts, we propose an in-context learning framework called PathGenIC that integrates context derived from the training set with a multimodal in-context learning (ICL) mechanism. Our method dynamically retrieves semantically similar whole slide image (WSI)-report pairs and incorporates adaptive feedback to enhance contextual relevance and generation quality. Evaluated on the HistGen benchmark, the framework achieves state-of-the-art results, with significant improvements across BLEU, METEOR, and ROUGE-L metrics, and demonstrates robustness across diverse report lengths and disease categories. By maximizing training data utility and bridging vision and language with ICL, our work offers a solution for AI-driven histopathology reporting, setting a strong foundation for future advancements in multimodal clinical applications. |
| title | Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2506.17645 |