OralGPT-Omni: A Versatile Dental Multimodal Large Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hao, Jing, Liang, Yuci, Lin, Lizhuo, Fan, Yuxuan, Zhou, Wenkai, Guo, Kaixin, Ye, Zanting, Sun, Yanpeng, Zhang, Xinyu, Yang, Yanqi, Li, Qiankun, Tang, Hao, Tsoi, James Kit-Hon, Shen, Linlin, Hung, Kuo Feng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915640410898432
author Hao, Jing
Liang, Yuci
Lin, Lizhuo
Fan, Yuxuan
Zhou, Wenkai
Guo, Kaixin
Ye, Zanting
Sun, Yanpeng
Zhang, Xinyu
Yang, Yanqi
Li, Qiankun
Tang, Hao
Tsoi, James Kit-Hon
Shen, Linlin
Hung, Kuo Feng
author_facet Hao, Jing
Liang, Yuci
Lin, Lizhuo
Fan, Yuxuan
Zhou, Wenkai
Guo, Kaixin
Ye, Zanting
Sun, Yanpeng
Zhang, Xinyu
Yang, Yanqi
Li, Qiankun
Tang, Hao
Tsoi, James Kit-Hon
Shen, Linlin
Hung, Kuo Feng
contents Multimodal Large Language Models (MLLMs) have exhibited immense potential across numerous medical specialties; yet, dentistry remains underexplored, in part due to limited domain-specific data, scarce dental expert annotations, insufficient modality-specific modeling, and challenges in reliability. In this paper, we present OralGPT-Omni, the first dental-specialized MLLM designed for comprehensive and trustworthy analysis across diverse dental imaging modalities and clinical tasks. To explicitly capture dentists' diagnostic reasoning, we construct TRACE-CoT, a clinically grounded chain-of-thought dataset that mirrors dental radiologists' decision-making processes. This reasoning supervision, combined with our proposed four-stage training paradigm, substantially strengthens the model's capacity for dental image understanding and analysis. In parallel, we introduce MMOral-Uni, the first unified multimodal benchmark for dental image analysis. It comprises 2,809 open-ended question-answer pairs spanning five modalities and five tasks, offering a comprehensive evaluation suite to date for MLLMs in digital dentistry. OralGPT-Omni achieves an overall score of 51.84 on the MMOral-Uni benchmark and 45.31 on the MMOral-OPG benchmark, dramatically outperforming the scores of GPT-5. Our work promotes intelligent dentistry and paves the way for future advances in dental image analysis. All code, benchmark, and models will be made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2511_22055
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OralGPT-Omni: A Versatile Dental Multimodal Large Language Model
Hao, Jing
Liang, Yuci
Lin, Lizhuo
Fan, Yuxuan
Zhou, Wenkai
Guo, Kaixin
Ye, Zanting
Sun, Yanpeng
Zhang, Xinyu
Yang, Yanqi
Li, Qiankun
Tang, Hao
Tsoi, James Kit-Hon
Shen, Linlin
Hung, Kuo Feng
Computer Vision and Pattern Recognition
Multimedia
Multimodal Large Language Models (MLLMs) have exhibited immense potential across numerous medical specialties; yet, dentistry remains underexplored, in part due to limited domain-specific data, scarce dental expert annotations, insufficient modality-specific modeling, and challenges in reliability. In this paper, we present OralGPT-Omni, the first dental-specialized MLLM designed for comprehensive and trustworthy analysis across diverse dental imaging modalities and clinical tasks. To explicitly capture dentists' diagnostic reasoning, we construct TRACE-CoT, a clinically grounded chain-of-thought dataset that mirrors dental radiologists' decision-making processes. This reasoning supervision, combined with our proposed four-stage training paradigm, substantially strengthens the model's capacity for dental image understanding and analysis. In parallel, we introduce MMOral-Uni, the first unified multimodal benchmark for dental image analysis. It comprises 2,809 open-ended question-answer pairs spanning five modalities and five tasks, offering a comprehensive evaluation suite to date for MLLMs in digital dentistry. OralGPT-Omni achieves an overall score of 51.84 on the MMOral-Uni benchmark and 45.31 on the MMOral-OPG benchmark, dramatically outperforming the scores of GPT-5. Our work promotes intelligent dentistry and paves the way for future advances in dental image analysis. All code, benchmark, and models will be made publicly available.
title OralGPT-Omni: A Versatile Dental Multimodal Large Language Model
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2511.22055