A benchmark multimodal oro-dental dataset for large vision-language models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lv, Haoxin, Haq, Ijazul, Du, Jin, Ma, Jiaxin, Zhu, Binnian, Dang, Xiaobing, Liang, Chaoan, Du, Ruxu, Zhang, Yingjie, Saqib, Muhammad
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918190245740544
author Lv, Haoxin
Haq, Ijazul
Du, Jin
Ma, Jiaxin
Zhu, Binnian
Dang, Xiaobing
Liang, Chaoan
Du, Ruxu
Zhang, Yingjie
Saqib, Muhammad
author_facet Lv, Haoxin
Haq, Ijazul
Du, Jin
Ma, Jiaxin
Zhu, Binnian
Dang, Xiaobing
Liang, Chaoan
Du, Ruxu
Zhang, Yingjie
Saqib, Muhammad
contents The advancement of artificial intelligence in oral healthcare relies on the availability of large-scale multimodal datasets that capture the complexity of clinical practice. In this paper, we present a comprehensive multimodal dataset, comprising 8775 dental checkups from 4800 patients collected over eight years (2018-2025), with patients ranging from 10 to 90 years of age. The dataset includes 50000 intraoral images, 8056 radiographs, and detailed textual records, including diagnoses, treatment plans, and follow-up notes. The data were collected under standard ethical guidelines and annotated for benchmarking. To demonstrate its utility, we fine-tuned state-of-the-art large vision-language models, Qwen-VL 3B and 7B, and evaluated them on two tasks: classification of six oro-dental anomalies and generation of complete diagnostic reports from multimodal inputs. We compared the fine-tuned models with their base counterparts and GPT-4o. The fine-tuned models achieved substantial gains over these baselines, validating the dataset and underscoring its effectiveness in advancing AI-driven oro-dental healthcare solutions. The dataset is publicly available, providing an essential resource for future research in AI dentistry.
format Preprint
id arxiv_https___arxiv_org_abs_2511_04948
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A benchmark multimodal oro-dental dataset for large vision-language models
Lv, Haoxin
Haq, Ijazul
Du, Jin
Ma, Jiaxin
Zhu, Binnian
Dang, Xiaobing
Liang, Chaoan
Du, Ruxu
Zhang, Yingjie
Saqib, Muhammad
Computer Vision and Pattern Recognition
Artificial Intelligence
The advancement of artificial intelligence in oral healthcare relies on the availability of large-scale multimodal datasets that capture the complexity of clinical practice. In this paper, we present a comprehensive multimodal dataset, comprising 8775 dental checkups from 4800 patients collected over eight years (2018-2025), with patients ranging from 10 to 90 years of age. The dataset includes 50000 intraoral images, 8056 radiographs, and detailed textual records, including diagnoses, treatment plans, and follow-up notes. The data were collected under standard ethical guidelines and annotated for benchmarking. To demonstrate its utility, we fine-tuned state-of-the-art large vision-language models, Qwen-VL 3B and 7B, and evaluated them on two tasks: classification of six oro-dental anomalies and generation of complete diagnostic reports from multimodal inputs. We compared the fine-tuned models with their base counterparts and GPT-4o. The fine-tuned models achieved substantial gains over these baselines, validating the dataset and underscoring its effectiveness in advancing AI-driven oro-dental healthcare solutions. The dataset is publicly available, providing an essential resource for future research in AI dentistry.
title A benchmark multimodal oro-dental dataset for large vision-language models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.04948