Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Nguyen, Huu Tien, Nguyen, Dac Thai, Nguyen, The Minh Duc, Nguyen, Trung Thanh, Truong, Thao Nguyen, Pham, Huy Hieu, Barthelemy, Johan, Tran, Minh Quan, Nguyen, Thanh Tam, Nguyen, Quoc Viet Hung, Chau, Quynh Anh, Mai, Hong Son, Nguyen, Thanh Trung, Nguyen, Phi Le
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917237636464640
author Nguyen, Huu Tien
Nguyen, Dac Thai
Nguyen, The Minh Duc
Nguyen, Trung Thanh
Truong, Thao Nguyen
Pham, Huy Hieu
Barthelemy, Johan
Tran, Minh Quan
Nguyen, Thanh Tam
Nguyen, Quoc Viet Hung
Chau, Quynh Anh
Mai, Hong Son
Nguyen, Thanh Trung
Nguyen, Phi Le
author_facet Nguyen, Huu Tien
Nguyen, Dac Thai
Nguyen, The Minh Duc
Nguyen, Trung Thanh
Truong, Thao Nguyen
Pham, Huy Hieu
Barthelemy, Johan
Tran, Minh Quan
Nguyen, Thanh Tam
Nguyen, Quoc Viet Hung
Chau, Quynh Anh
Mai, Hong Son
Nguyen, Thanh Trung
Nguyen, Phi Le
contents Vision-Language Foundation Models (VLMs), trained on large-scale multimodal datasets, have driven significant advances in Artificial Intelligence (AI) by enabling rich cross-modal reasoning. Despite their success in general domains, applying these models to medical imaging remains challenging due to the limited availability of diverse imaging modalities and multilingual clinical data. Most existing medical VLMs are trained on a subset of imaging modalities and focus primarily on high-resource languages, thus limiting their generalizability and clinical utility. To address these limitations, we introduce a novel Vietnamese-language multimodal medical dataset consisting of 2,757 whole-body PET/CT volumes from independent patients and their corresponding full-length clinical reports. This dataset is designed to fill two pressing gaps in medical AI development: (1) the lack of PET/CT imaging data in existing VLMs training corpora, which hinders the development of models capable of handling functional imaging tasks; and (2) the underrepresentation of low-resource languages, particularly the Vietnamese language, in medical vision-language research. To the best of our knowledge, this is the first dataset to provide comprehensive PET/CT-report pairs in Vietnamese. We further introduce a training framework to enhance VLMs' learning, including data augmentation and expert-validated test sets. We conduct comprehensive experiments benchmarking state-of-the-art VLMs on downstream tasks. The experimental results show that incorporating our dataset significantly improves the performance of existing VLMs. We believe this dataset and benchmark will serve as a pivotal step in advancing the development of more robust VLMs for medical imaging, especially for low-resource languages and clinical use in Vietnamese healthcare. The source code is available at https://github.com/AIoT-Lab-BKAI/ViPET-ReportGen.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24739
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation
Nguyen, Huu Tien
Nguyen, Dac Thai
Nguyen, The Minh Duc
Nguyen, Trung Thanh
Truong, Thao Nguyen
Pham, Huy Hieu
Barthelemy, Johan
Tran, Minh Quan
Nguyen, Thanh Tam
Nguyen, Quoc Viet Hung
Chau, Quynh Anh
Mai, Hong Son
Nguyen, Thanh Trung
Nguyen, Phi Le
Computer Vision and Pattern Recognition
Vision-Language Foundation Models (VLMs), trained on large-scale multimodal datasets, have driven significant advances in Artificial Intelligence (AI) by enabling rich cross-modal reasoning. Despite their success in general domains, applying these models to medical imaging remains challenging due to the limited availability of diverse imaging modalities and multilingual clinical data. Most existing medical VLMs are trained on a subset of imaging modalities and focus primarily on high-resource languages, thus limiting their generalizability and clinical utility. To address these limitations, we introduce a novel Vietnamese-language multimodal medical dataset consisting of 2,757 whole-body PET/CT volumes from independent patients and their corresponding full-length clinical reports. This dataset is designed to fill two pressing gaps in medical AI development: (1) the lack of PET/CT imaging data in existing VLMs training corpora, which hinders the development of models capable of handling functional imaging tasks; and (2) the underrepresentation of low-resource languages, particularly the Vietnamese language, in medical vision-language research. To the best of our knowledge, this is the first dataset to provide comprehensive PET/CT-report pairs in Vietnamese. We further introduce a training framework to enhance VLMs' learning, including data augmentation and expert-validated test sets. We conduct comprehensive experiments benchmarking state-of-the-art VLMs on downstream tasks. The experimental results show that incorporating our dataset significantly improves the performance of existing VLMs. We believe this dataset and benchmark will serve as a pivotal step in advancing the development of more robust VLMs for medical imaging, especially for low-resource languages and clinical use in Vietnamese healthcare. The source code is available at https://github.com/AIoT-Lab-BKAI/ViPET-ReportGen.
title Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.24739