Boosting Medical Image-based Cancer Detection via Text-guided Supervision from Reports

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Guo, Guangyu, Yao, Jiawen, Xia, Yingda, Mok, Tony C. W., Zheng, Zhilin, Han, Junwei, Lu, Le, Zhang, Dingwen, Zhou, Jian, Zhang, Ling
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929354683973632
author Guo, Guangyu
Yao, Jiawen
Xia, Yingda
Mok, Tony C. W.
Zheng, Zhilin
Han, Junwei
Lu, Le
Zhang, Dingwen
Zhou, Jian
Zhang, Ling
author_facet Guo, Guangyu
Yao, Jiawen
Xia, Yingda
Mok, Tony C. W.
Zheng, Zhilin
Han, Junwei
Lu, Le
Zhang, Dingwen
Zhou, Jian
Zhang, Ling
contents The absence of adequately sufficient expert-level tumor annotations hinders the effectiveness of supervised learning based opportunistic cancer screening on medical imaging. Clinical reports (that are rich in descriptive textual details) can offer a "free lunch'' supervision information and provide tumor location as a type of weak label to cope with screening tasks, thus saving human labeling workloads, if properly leveraged. However, predicting cancer only using such weak labels can be very changeling since tumors are usually presented in small anatomical regions compared to the whole 3D medical scans. Weakly semi-supervised learning (WSSL) utilizes a limited set of voxel-level tumor annotations and incorporates alongside a substantial number of medical images that have only off-the-shelf clinical reports, which may strike a good balance between minimizing expert annotation workload and optimizing screening efficacy. In this paper, we propose a novel text-guided learning method to achieve highly accurate cancer detection results. Through integrating diagnostic and tumor location text prompts into the text encoder of a vision-language model (VLM), optimization of weakly supervised learning can be effectively performed in the latent space of VLM, thereby enhancing the stability of training. Our approach can leverage clinical knowledge by large-scale pre-trained VLM to enhance generalization ability, and produce reliable pseudo tumor masks to improve cancer detection. Our extensive quantitative experimental results on a large-scale cancer dataset, including 1,651 unique patients, validate that our approach can reduce human annotation efforts by at least 70% while maintaining comparable cancer detection accuracy to competing fully supervised methods (AUC value 0.961 versus 0.966).
format Preprint
id arxiv_https___arxiv_org_abs_2405_14230
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Boosting Medical Image-based Cancer Detection via Text-guided Supervision from Reports
Guo, Guangyu
Yao, Jiawen
Xia, Yingda
Mok, Tony C. W.
Zheng, Zhilin
Han, Junwei
Lu, Le
Zhang, Dingwen
Zhou, Jian
Zhang, Ling
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
The absence of adequately sufficient expert-level tumor annotations hinders the effectiveness of supervised learning based opportunistic cancer screening on medical imaging. Clinical reports (that are rich in descriptive textual details) can offer a "free lunch'' supervision information and provide tumor location as a type of weak label to cope with screening tasks, thus saving human labeling workloads, if properly leveraged. However, predicting cancer only using such weak labels can be very changeling since tumors are usually presented in small anatomical regions compared to the whole 3D medical scans. Weakly semi-supervised learning (WSSL) utilizes a limited set of voxel-level tumor annotations and incorporates alongside a substantial number of medical images that have only off-the-shelf clinical reports, which may strike a good balance between minimizing expert annotation workload and optimizing screening efficacy. In this paper, we propose a novel text-guided learning method to achieve highly accurate cancer detection results. Through integrating diagnostic and tumor location text prompts into the text encoder of a vision-language model (VLM), optimization of weakly supervised learning can be effectively performed in the latent space of VLM, thereby enhancing the stability of training. Our approach can leverage clinical knowledge by large-scale pre-trained VLM to enhance generalization ability, and produce reliable pseudo tumor masks to improve cancer detection. Our extensive quantitative experimental results on a large-scale cancer dataset, including 1,651 unique patients, validate that our approach can reduce human annotation efforts by at least 70% while maintaining comparable cancer detection accuracy to competing fully supervised methods (AUC value 0.961 versus 0.966).
title Boosting Medical Image-based Cancer Detection via Text-guided Supervision from Reports
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2405.14230