OralGPT-Omni: A Versatile Dental Multimodal Large Language Model
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915640410898432 |
|---|---|
| author | Hao, Jing Liang, Yuci Lin, Lizhuo Fan, Yuxuan Zhou, Wenkai Guo, Kaixin Ye, Zanting Sun, Yanpeng Zhang, Xinyu Yang, Yanqi Li, Qiankun Tang, Hao Tsoi, James Kit-Hon Shen, Linlin Hung, Kuo Feng |
| author_facet | Hao, Jing Liang, Yuci Lin, Lizhuo Fan, Yuxuan Zhou, Wenkai Guo, Kaixin Ye, Zanting Sun, Yanpeng Zhang, Xinyu Yang, Yanqi Li, Qiankun Tang, Hao Tsoi, James Kit-Hon Shen, Linlin Hung, Kuo Feng |
| contents | Multimodal Large Language Models (MLLMs) have exhibited immense potential across numerous medical specialties; yet, dentistry remains underexplored, in part due to limited domain-specific data, scarce dental expert annotations, insufficient modality-specific modeling, and challenges in reliability. In this paper, we present OralGPT-Omni, the first dental-specialized MLLM designed for comprehensive and trustworthy analysis across diverse dental imaging modalities and clinical tasks. To explicitly capture dentists' diagnostic reasoning, we construct TRACE-CoT, a clinically grounded chain-of-thought dataset that mirrors dental radiologists' decision-making processes. This reasoning supervision, combined with our proposed four-stage training paradigm, substantially strengthens the model's capacity for dental image understanding and analysis. In parallel, we introduce MMOral-Uni, the first unified multimodal benchmark for dental image analysis. It comprises 2,809 open-ended question-answer pairs spanning five modalities and five tasks, offering a comprehensive evaluation suite to date for MLLMs in digital dentistry. OralGPT-Omni achieves an overall score of 51.84 on the MMOral-Uni benchmark and 45.31 on the MMOral-OPG benchmark, dramatically outperforming the scores of GPT-5. Our work promotes intelligent dentistry and paves the way for future advances in dental image analysis. All code, benchmark, and models will be made publicly available. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_22055 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | OralGPT-Omni: A Versatile Dental Multimodal Large Language Model Hao, Jing Liang, Yuci Lin, Lizhuo Fan, Yuxuan Zhou, Wenkai Guo, Kaixin Ye, Zanting Sun, Yanpeng Zhang, Xinyu Yang, Yanqi Li, Qiankun Tang, Hao Tsoi, James Kit-Hon Shen, Linlin Hung, Kuo Feng Computer Vision and Pattern Recognition Multimedia Multimodal Large Language Models (MLLMs) have exhibited immense potential across numerous medical specialties; yet, dentistry remains underexplored, in part due to limited domain-specific data, scarce dental expert annotations, insufficient modality-specific modeling, and challenges in reliability. In this paper, we present OralGPT-Omni, the first dental-specialized MLLM designed for comprehensive and trustworthy analysis across diverse dental imaging modalities and clinical tasks. To explicitly capture dentists' diagnostic reasoning, we construct TRACE-CoT, a clinically grounded chain-of-thought dataset that mirrors dental radiologists' decision-making processes. This reasoning supervision, combined with our proposed four-stage training paradigm, substantially strengthens the model's capacity for dental image understanding and analysis. In parallel, we introduce MMOral-Uni, the first unified multimodal benchmark for dental image analysis. It comprises 2,809 open-ended question-answer pairs spanning five modalities and five tasks, offering a comprehensive evaluation suite to date for MLLMs in digital dentistry. OralGPT-Omni achieves an overall score of 51.84 on the MMOral-Uni benchmark and 45.31 on the MMOral-OPG benchmark, dramatically outperforming the scores of GPT-5. Our work promotes intelligent dentistry and paves the way for future advances in dental image analysis. All code, benchmark, and models will be made publicly available. |
| title | OralGPT-Omni: A Versatile Dental Multimodal Large Language Model |
| topic | Computer Vision and Pattern Recognition Multimedia |
| url | https://arxiv.org/abs/2511.22055 |