Cost-effective Instruction Learning for Pathology Vision and Language Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Kaitao, Liu, Mianxin, Yan, Fang, Ma, Lei, Shi, Xiaoming, Wang, Lilong, Wang, Xiaosong, Zhu, Lifeng, Wang, Zhe, Zhou, Mu, Zhang, Shaoting
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908414042439680
author Chen, Kaitao
Liu, Mianxin
Yan, Fang
Ma, Lei
Shi, Xiaoming
Wang, Lilong
Wang, Xiaosong
Zhu, Lifeng
Wang, Zhe
Zhou, Mu
Zhang, Shaoting
author_facet Chen, Kaitao
Liu, Mianxin
Yan, Fang
Ma, Lei
Shi, Xiaoming
Wang, Lilong
Wang, Xiaosong
Zhu, Lifeng
Wang, Zhe
Zhou, Mu
Zhang, Shaoting
contents The advent of vision-language models fosters the interactive conversations between AI-enabled models and humans. Yet applying these models into clinics must deal with daunting challenges around large-scale training data, financial, and computational resources. Here we propose a cost-effective instruction learning framework for conversational pathology named as CLOVER. CLOVER only trains a lightweight module and uses instruction tuning while freezing the parameters of the large language model. Instead of using costly GPT-4, we propose well-designed prompts on GPT-3.5 for building generation-based instructions, emphasizing the utility of pathological knowledge derived from the Internet source. To augment the use of instructions, we construct a high-quality set of template-based instructions in the context of digital pathology. From two benchmark datasets, our findings reveal the strength of hybrid-form instructions in the visual question-answer in pathology. Extensive results show the cost-effectiveness of CLOVER in answering both open-ended and closed-ended questions, where CLOVER outperforms strong baselines that possess 37 times more training parameters and use instruction data generated from GPT-4. Through the instruction tuning, CLOVER exhibits robustness of few-shot learning in the external clinical dataset. These findings demonstrate that cost-effective modeling of CLOVER could accelerate the adoption of rapid conversational applications in the landscape of digital pathology.
format Preprint
id arxiv_https___arxiv_org_abs_2407_17734
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Cost-effective Instruction Learning for Pathology Vision and Language Analysis
Chen, Kaitao
Liu, Mianxin
Yan, Fang
Ma, Lei
Shi, Xiaoming
Wang, Lilong
Wang, Xiaosong
Zhu, Lifeng
Wang, Zhe
Zhou, Mu
Zhang, Shaoting
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
The advent of vision-language models fosters the interactive conversations between AI-enabled models and humans. Yet applying these models into clinics must deal with daunting challenges around large-scale training data, financial, and computational resources. Here we propose a cost-effective instruction learning framework for conversational pathology named as CLOVER. CLOVER only trains a lightweight module and uses instruction tuning while freezing the parameters of the large language model. Instead of using costly GPT-4, we propose well-designed prompts on GPT-3.5 for building generation-based instructions, emphasizing the utility of pathological knowledge derived from the Internet source. To augment the use of instructions, we construct a high-quality set of template-based instructions in the context of digital pathology. From two benchmark datasets, our findings reveal the strength of hybrid-form instructions in the visual question-answer in pathology. Extensive results show the cost-effectiveness of CLOVER in answering both open-ended and closed-ended questions, where CLOVER outperforms strong baselines that possess 37 times more training parameters and use instruction data generated from GPT-4. Through the instruction tuning, CLOVER exhibits robustness of few-shot learning in the external clinical dataset. These findings demonstrate that cost-effective modeling of CLOVER could accelerate the adoption of rapid conversational applications in the landscape of digital pathology.
title Cost-effective Instruction Learning for Pathology Vision and Language Analysis
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.17734