DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhu, Hengchuan, Xu, Yihuan, Li, Yichen, Meng, Zijie, Liu, Zuozhu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918131905069056
author Zhu, Hengchuan
Xu, Yihuan
Li, Yichen
Meng, Zijie
Liu, Zuozhu
author_facet Zhu, Hengchuan
Xu, Yihuan
Li, Yichen
Meng, Zijie
Liu, Zuozhu
contents Recent advances in large language models (LLMs) and medical LLMs (Med-LLMs) have demonstrated strong performance on general medical benchmarks. However, their capabilities in specialized medical fields, such as dentistry which require deeper domain-specific knowledge, remain underexplored due to the lack of targeted evaluation resources. In this paper, we introduce DentalBench, the first comprehensive bilingual benchmark designed to evaluate and advance LLMs in the dental domain. DentalBench consists of two main components: DentalQA, an English-Chinese question-answering (QA) benchmark with 36,597 questions spanning 4 tasks and 16 dental subfields; and DentalCorpus, a large-scale, high-quality corpus with 337.35 million tokens curated for dental domain adaptation, supporting both supervised fine-tuning (SFT) and retrieval-augmented generation (RAG). We evaluate 14 LLMs, covering proprietary, open-source, and medical-specific models, and reveal significant performance gaps across task types and languages. Further experiments with Qwen-2.5-3B demonstrate that domain adaptation substantially improves model performance, particularly on knowledge-intensive and terminology-focused tasks, and highlight the importance of domain-specific benchmarks for developing trustworthy and effective LLMs tailored to healthcare applications.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20416
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
Zhu, Hengchuan
Xu, Yihuan
Li, Yichen
Meng, Zijie
Liu, Zuozhu
Computation and Language
Artificial Intelligence
Recent advances in large language models (LLMs) and medical LLMs (Med-LLMs) have demonstrated strong performance on general medical benchmarks. However, their capabilities in specialized medical fields, such as dentistry which require deeper domain-specific knowledge, remain underexplored due to the lack of targeted evaluation resources. In this paper, we introduce DentalBench, the first comprehensive bilingual benchmark designed to evaluate and advance LLMs in the dental domain. DentalBench consists of two main components: DentalQA, an English-Chinese question-answering (QA) benchmark with 36,597 questions spanning 4 tasks and 16 dental subfields; and DentalCorpus, a large-scale, high-quality corpus with 337.35 million tokens curated for dental domain adaptation, supporting both supervised fine-tuning (SFT) and retrieval-augmented generation (RAG). We evaluate 14 LLMs, covering proprietary, open-source, and medical-specific models, and reveal significant performance gaps across task types and languages. Further experiments with Qwen-2.5-3B demonstrate that domain adaptation substantially improves model performance, particularly on knowledge-intensive and terminology-focused tasks, and highlight the importance of domain-specific benchmarks for developing trustworthy and effective LLMs tailored to healthcare applications.
title DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.20416