E3D-GPT: Enhanced 3D Visual Foundation for Medical Vision-Language Model

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lai, Haoran, Jiang, Zihang, Yao, Qingsong, Wang, Rongsheng, He, Zhiyang, Tao, Xiaodong, Wei, Wei, Lv, Weifu, Zhou, S. Kevin
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929549771538432
author Lai, Haoran
Jiang, Zihang
Yao, Qingsong
Wang, Rongsheng
He, Zhiyang
Tao, Xiaodong
Wei, Wei
Lv, Weifu
Zhou, S. Kevin
author_facet Lai, Haoran
Jiang, Zihang
Yao, Qingsong
Wang, Rongsheng
He, Zhiyang
Tao, Xiaodong
Wei, Wei
Lv, Weifu
Zhou, S. Kevin
contents The development of 3D medical vision-language models holds significant potential for disease diagnosis and patient treatment. However, compared to 2D medical images, 3D medical images, such as CT scans, face challenges related to limited training data and high dimension, which severely restrict the progress of 3D medical vision-language models. To address these issues, we collect a large amount of unlabeled 3D CT data and utilize self-supervised learning to construct a 3D visual foundation model for extracting 3D visual features. Then, we apply 3D spatial convolutions to aggregate and project high-level image features, reducing computational complexity while preserving spatial information. We also construct two instruction-tuning datasets based on BIMCV-R and CT-RATE to fine-tune the 3D vision-language model. Our model demonstrates superior performance compared to existing methods in report generation, visual question answering, and disease diagnosis. Code and data will be made publicly available soon.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14200
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle E3D-GPT: Enhanced 3D Visual Foundation for Medical Vision-Language Model
Lai, Haoran
Jiang, Zihang
Yao, Qingsong
Wang, Rongsheng
He, Zhiyang
Tao, Xiaodong
Wei, Wei
Lv, Weifu
Zhou, S. Kevin
Image and Video Processing
Computation and Language
Computer Vision and Pattern Recognition
The development of 3D medical vision-language models holds significant potential for disease diagnosis and patient treatment. However, compared to 2D medical images, 3D medical images, such as CT scans, face challenges related to limited training data and high dimension, which severely restrict the progress of 3D medical vision-language models. To address these issues, we collect a large amount of unlabeled 3D CT data and utilize self-supervised learning to construct a 3D visual foundation model for extracting 3D visual features. Then, we apply 3D spatial convolutions to aggregate and project high-level image features, reducing computational complexity while preserving spatial information. We also construct two instruction-tuning datasets based on BIMCV-R and CT-RATE to fine-tune the 3D vision-language model. Our model demonstrates superior performance compared to existing methods in report generation, visual question answering, and disease diagnosis. Code and data will be made publicly available soon.
title E3D-GPT: Enhanced 3D Visual Foundation for Medical Vision-Language Model
topic Image and Video Processing
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.14200