Enhancing the vision-language foundation model with key semantic knowledge-emphasized report refinement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Weijian, Li, Cheng, Yang, Hao, Liu, Jiarun, Liang, Yong, Zheng, Hairong, Wang, Shanshan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916381407051776
author Huang, Weijian
Li, Cheng
Yang, Hao
Liu, Jiarun
Liang, Yong
Zheng, Hairong
Wang, Shanshan
author_facet Huang, Weijian
Li, Cheng
Yang, Hao
Liu, Jiarun
Liang, Yong
Zheng, Hairong
Wang, Shanshan
contents Recently, vision-language representation learning has made remarkable advancements in building up medical foundation models, holding immense potential for transforming the landscape of clinical research and medical care. The underlying hypothesis is that the rich knowledge embedded in radiology reports can effectively assist and guide the learning process, reducing the need for additional labels. However, these reports tend to be complex and sometimes even consist of redundant descriptions that make the representation learning too challenging to capture the key semantic information. This paper develops a novel iterative vision-language representation learning framework by proposing a key semantic knowledge-emphasized report refinement method. Particularly, raw radiology reports are refined to highlight the key information according to a constructed clinical dictionary and two model-optimized knowledge-enhancement metrics. The iterative framework is designed to progressively learn, starting from gaining a general understanding of the patient's condition based on raw reports and gradually refines and extracts critical information essential to the fine-grained analysis tasks. The effectiveness of the proposed framework is validated on various downstream medical image analysis tasks, including disease classification, region-of-interest segmentation, and phrase grounding. Our framework surpasses seven state-of-the-art methods in both fine-tuning and zero-shot settings, demonstrating its encouraging potential for different clinical applications.
format Preprint
id arxiv_https___arxiv_org_abs_2401_11421
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing the vision-language foundation model with key semantic knowledge-emphasized report refinement
Huang, Weijian
Li, Cheng
Yang, Hao
Liu, Jiarun
Liang, Yong
Zheng, Hairong
Wang, Shanshan
Computer Vision and Pattern Recognition
Recently, vision-language representation learning has made remarkable advancements in building up medical foundation models, holding immense potential for transforming the landscape of clinical research and medical care. The underlying hypothesis is that the rich knowledge embedded in radiology reports can effectively assist and guide the learning process, reducing the need for additional labels. However, these reports tend to be complex and sometimes even consist of redundant descriptions that make the representation learning too challenging to capture the key semantic information. This paper develops a novel iterative vision-language representation learning framework by proposing a key semantic knowledge-emphasized report refinement method. Particularly, raw radiology reports are refined to highlight the key information according to a constructed clinical dictionary and two model-optimized knowledge-enhancement metrics. The iterative framework is designed to progressively learn, starting from gaining a general understanding of the patient's condition based on raw reports and gradually refines and extracts critical information essential to the fine-grained analysis tasks. The effectiveness of the proposed framework is validated on various downstream medical image analysis tasks, including disease classification, region-of-interest segmentation, and phrase grounding. Our framework surpasses seven state-of-the-art methods in both fine-tuning and zero-shot settings, demonstrating its encouraging potential for different clinical applications.
title Enhancing the vision-language foundation model with key semantic knowledge-emphasized report refinement
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.11421