ViX-Ray: A Vietnamese Chest X-Ray Dataset for Vision-Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Nguyen, Duy Vu Minh, Truong, Chinh Thanh, Tran, Phuc Hoang, Le, Hung Tuan, Dat, Nguyen Van-Thanh, Pham, Trung Hieu, Van Nguyen, Kiet
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911520102809600
author Nguyen, Duy Vu Minh
Truong, Chinh Thanh
Tran, Phuc Hoang
Le, Hung Tuan
Dat, Nguyen Van-Thanh
Pham, Trung Hieu
Van Nguyen, Kiet
author_facet Nguyen, Duy Vu Minh
Truong, Chinh Thanh
Tran, Phuc Hoang
Le, Hung Tuan
Dat, Nguyen Van-Thanh
Pham, Trung Hieu
Van Nguyen, Kiet
contents Vietnamese medical research has become an increasingly vital domain, particularly with the rise of intelligent technologies aimed at reducing time and resource burdens in clinical diagnosis. Recent advances in vision-language models (VLMs), such as Gemini and GPT-4V, have sparked a growing interest in applying AI to healthcare. However, most existing VLMs lack exposure to Vietnamese medical data, limiting their ability to generate accurate and contextually appropriate diagnostic outputs for Vietnamese patients. To address this challenge, we introduce ViX-Ray, a novel dataset comprising 5,400 Vietnamese chest X-ray images annotated with expert-written findings and impressions from physicians at a major Vietnamese hospital. We analyze linguistic patterns within the dataset, including the frequency of mentioned body parts and diagnoses, to identify domain-specific linguistic characteristics of Vietnamese radiology reports. Furthermore, we fine-tune five state-of-the-art open-source VLMs on ViX-Ray and compare their performance to leading proprietary models, GPT-4V and Gemini. Our results show that while several models generate outputs partially aligned with clinical ground truths, they often suffer from low precision and excessive hallucination, especially in impression generation. These findings not only demonstrate the complexity and challenge of our dataset but also establish ViX-Ray as a valuable benchmark for evaluating and advancing vision-language models in the Vietnamese clinical domain.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15513
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ViX-Ray: A Vietnamese Chest X-Ray Dataset for Vision-Language Models
Nguyen, Duy Vu Minh
Truong, Chinh Thanh
Tran, Phuc Hoang
Le, Hung Tuan
Dat, Nguyen Van-Thanh
Pham, Trung Hieu
Van Nguyen, Kiet
Computation and Language
Vietnamese medical research has become an increasingly vital domain, particularly with the rise of intelligent technologies aimed at reducing time and resource burdens in clinical diagnosis. Recent advances in vision-language models (VLMs), such as Gemini and GPT-4V, have sparked a growing interest in applying AI to healthcare. However, most existing VLMs lack exposure to Vietnamese medical data, limiting their ability to generate accurate and contextually appropriate diagnostic outputs for Vietnamese patients. To address this challenge, we introduce ViX-Ray, a novel dataset comprising 5,400 Vietnamese chest X-ray images annotated with expert-written findings and impressions from physicians at a major Vietnamese hospital. We analyze linguistic patterns within the dataset, including the frequency of mentioned body parts and diagnoses, to identify domain-specific linguistic characteristics of Vietnamese radiology reports. Furthermore, we fine-tune five state-of-the-art open-source VLMs on ViX-Ray and compare their performance to leading proprietary models, GPT-4V and Gemini. Our results show that while several models generate outputs partially aligned with clinical ground truths, they often suffer from low precision and excessive hallucination, especially in impression generation. These findings not only demonstrate the complexity and challenge of our dataset but also establish ViX-Ray as a valuable benchmark for evaluating and advancing vision-language models in the Vietnamese clinical domain.
title ViX-Ray: A Vietnamese Chest X-Ray Dataset for Vision-Language Models
topic Computation and Language
url https://arxiv.org/abs/2603.15513