Discovering Pathology Rationale and Token Allocation for Efficient Multimodal Pathology Reasoning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xu, Zhe, Jin, Cheng, Wang, Yihui, Liu, Ziyi, Chen, Hao
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908373497151488
author Xu, Zhe
Jin, Cheng
Wang, Yihui
Liu, Ziyi
Chen, Hao
author_facet Xu, Zhe
Jin, Cheng
Wang, Yihui
Liu, Ziyi
Chen, Hao
contents Multimodal pathological image understanding has garnered widespread interest due to its potential to improve diagnostic accuracy and enable personalized treatment through integrated visual and textual data. However, existing methods exhibit limited reasoning capabilities, which hamper their ability to handle complex diagnostic scenarios. Additionally, the enormous size of pathological images leads to severe computational burdens, further restricting their practical deployment. To address these limitations, we introduce a novel bilateral reinforcement learning framework comprising two synergistic branches. One reinforcement branch enhances the reasoning capability by enabling the model to learn task-specific decision processes, i.e., pathology rationales, directly from labels without explicit reasoning supervision. While the other branch dynamically allocates a tailored number of tokens to different images based on both their visual content and task context, thereby optimizing computational efficiency. We apply our method to various pathological tasks such as visual question answering, cancer subtyping, and lesion detection. Extensive experiments show an average +41.7 absolute performance improvement with 70.3% lower inference costs over the base models, achieving both reasoning accuracy and computational efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15687
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Discovering Pathology Rationale and Token Allocation for Efficient Multimodal Pathology Reasoning
Xu, Zhe
Jin, Cheng
Wang, Yihui
Liu, Ziyi
Chen, Hao
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimodal pathological image understanding has garnered widespread interest due to its potential to improve diagnostic accuracy and enable personalized treatment through integrated visual and textual data. However, existing methods exhibit limited reasoning capabilities, which hamper their ability to handle complex diagnostic scenarios. Additionally, the enormous size of pathological images leads to severe computational burdens, further restricting their practical deployment. To address these limitations, we introduce a novel bilateral reinforcement learning framework comprising two synergistic branches. One reinforcement branch enhances the reasoning capability by enabling the model to learn task-specific decision processes, i.e., pathology rationales, directly from labels without explicit reasoning supervision. While the other branch dynamically allocates a tailored number of tokens to different images based on both their visual content and task context, thereby optimizing computational efficiency. We apply our method to various pathological tasks such as visual question answering, cancer subtyping, and lesion detection. Extensive experiments show an average +41.7 absolute performance improvement with 70.3% lower inference costs over the base models, achieving both reasoning accuracy and computational efficiency.
title Discovering Pathology Rationale and Token Allocation for Efficient Multimodal Pathology Reasoning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.15687