OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Rongyang, Zhou, Shuang, Wang, Jiashuo, Xie, Wenya, Che, Xiaoxia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
von: Ge, Wentao, et al.
Veröffentlicht: (2023)
von: Ge, Wentao, et al.
Veröffentlicht: (2023)
PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain
von: Chen, Liang, et al.
Veröffentlicht: (2024)
von: Chen, Liang, et al.
Veröffentlicht: (2024)
HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models
von: Que, Haoran, et al.
Veröffentlicht: (2024)
von: Que, Haoran, et al.
Veröffentlicht: (2024)
MLLM-CL: Continual Learning for Multimodal Large Language Models
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues
von: Kuo, Tzu-Lin, et al.
Veröffentlicht: (2024)
von: Kuo, Tzu-Lin, et al.
Veröffentlicht: (2024)
Evaluating Accounting Reasoning Capabilities of Large Language Models
von: Zhou, Jie, et al.
Veröffentlicht: (2026)
von: Zhou, Jie, et al.
Veröffentlicht: (2026)
DebugBench: Evaluating Debugging Capability of Large Language Models
von: Tian, Runchu, et al.
Veröffentlicht: (2024)
von: Tian, Runchu, et al.
Veröffentlicht: (2024)
CBT-Bench: Evaluating Large Language Models on Assisting Cognitive Behavior Therapy
von: Zhang, Mian, et al.
Veröffentlicht: (2024)
von: Zhang, Mian, et al.
Veröffentlicht: (2024)
Large Language Model-Enhanced Symbolic Reasoning for Knowledge Base Completion
von: He, Qiyuan, et al.
Veröffentlicht: (2025)
von: He, Qiyuan, et al.
Veröffentlicht: (2025)
FoundaBench: Evaluating Chinese Fundamental Knowledge Capabilities of Large Language Models
von: Li, Wei, et al.
Veröffentlicht: (2024)
von: Li, Wei, et al.
Veröffentlicht: (2024)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
CityBench: Evaluating the Capabilities of Large Language Models for Urban Tasks
von: Feng, Jie, et al.
Veröffentlicht: (2024)
von: Feng, Jie, et al.
Veröffentlicht: (2024)
Decoupled Proxy Alignment: Mitigating Language Prior Conflict for Multimodal Alignment in MLLM
von: Tan, Chenkun, et al.
Veröffentlicht: (2025)
von: Tan, Chenkun, et al.
Veröffentlicht: (2025)
TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators
von: Li, Jianling, et al.
Veröffentlicht: (2025)
von: Li, Jianling, et al.
Veröffentlicht: (2025)
Interpretable Differential Diagnosis with Dual-Inference Large Language Models
von: Zhou, Shuang, et al.
Veröffentlicht: (2024)
von: Zhou, Shuang, et al.
Veröffentlicht: (2024)
SNS-Bench-VL: Benchmarking Multimodal Large Language Models in Social Networking Services
von: Guo, Hongcheng, et al.
Veröffentlicht: (2025)
von: Guo, Hongcheng, et al.
Veröffentlicht: (2025)
Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models
von: Zeng, Bo, et al.
Veröffentlicht: (2025)
von: Zeng, Bo, et al.
Veröffentlicht: (2025)
FAC$^2$E: Better Understanding Large Language Model Capabilities by Dissociating Language and Cognition
von: Wang, Xiaoqiang, et al.
Veröffentlicht: (2024)
von: Wang, Xiaoqiang, et al.
Veröffentlicht: (2024)
RefineBench: Evaluating Refinement Capability of Language Models via Checklists
von: Lee, Young-Jun, et al.
Veröffentlicht: (2025)
von: Lee, Young-Jun, et al.
Veröffentlicht: (2025)
MLLM-LLaVA-FL: Multimodal Large Language Model Assisted Federated Learning
von: Zhang, Jianyi, et al.
Veröffentlicht: (2024)
von: Zhang, Jianyi, et al.
Veröffentlicht: (2024)
MentalBench: A DSM-Grounded Benchmark for Evaluating Psychiatric Diagnostic Capability of Large Language Models
von: Song, Hoyun, et al.
Veröffentlicht: (2026)
von: Song, Hoyun, et al.
Veröffentlicht: (2026)
CodeEditorBench: Evaluating Code Editing Capability of Large Language Models
von: Guo, Jiawei, et al.
Veröffentlicht: (2024)
von: Guo, Jiawei, et al.
Veröffentlicht: (2024)
BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems
von: Wang, Wei, et al.
Veröffentlicht: (2024)
von: Wang, Wei, et al.
Veröffentlicht: (2024)
IAA: Inner-Adaptor Architecture Empowers Frozen Large Language Model with Multimodal Capabilities
von: Wang, Bin, et al.
Veröffentlicht: (2024)
von: Wang, Bin, et al.
Veröffentlicht: (2024)
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages
von: Han, Wenhan, et al.
Veröffentlicht: (2025)
von: Han, Wenhan, et al.
Veröffentlicht: (2025)
Large Language Models are Capable of Offering Cognitive Reappraisal, if Guided
von: Zhan, Hongli, et al.
Veröffentlicht: (2024)
von: Zhan, Hongli, et al.
Veröffentlicht: (2024)
Can Language Models Act as Knowledge Bases at Scale?
von: He, Qiyuan, et al.
Veröffentlicht: (2024)
von: He, Qiyuan, et al.
Veröffentlicht: (2024)
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
von: Yu, Weihao, et al.
Veröffentlicht: (2023)
von: Yu, Weihao, et al.
Veröffentlicht: (2023)
Large Language Models Know What Makes Exemplary Contexts
von: Long, Quanyu, et al.
Veröffentlicht: (2024)
von: Long, Quanyu, et al.
Veröffentlicht: (2024)
PromptBench: A Unified Library for Evaluation of Large Language Models
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
NewsBench: A Systematic Evaluation Framework for Assessing Editorial Capabilities of Large Language Models in Chinese Journalism
von: Li, Miao, et al.
Veröffentlicht: (2024)
von: Li, Miao, et al.
Veröffentlicht: (2024)
MLLM-Microscope: Unlocking Hidden Structure Within Multimodal Large Language Models
von: Mussabayev, Ravil, et al.
Veröffentlicht: (2026)
von: Mussabayev, Ravil, et al.
Veröffentlicht: (2026)
MDIT-Bench: Evaluating the Dual-Implicit Toxicity in Large Multimodal Models
von: Jin, Bohan, et al.
Veröffentlicht: (2025)
von: Jin, Bohan, et al.
Veröffentlicht: (2025)
OphthBench: A Comprehensive Benchmark for Evaluating Large Language Models in Chinese Ophthalmology
von: Zhou, Chengfeng, et al.
Veröffentlicht: (2025)
von: Zhou, Chengfeng, et al.
Veröffentlicht: (2025)
EmoBench-M: Benchmarking Emotional Intelligence for Multimodal Large Language Models
von: Hu, He, et al.
Veröffentlicht: (2025)
von: Hu, He, et al.
Veröffentlicht: (2025)
FEEL: A Framework for Evaluating Emotional Support Capability with Large Language Models
von: Zhang, Huaiwen, et al.
Veröffentlicht: (2024)
von: Zhang, Huaiwen, et al.
Veröffentlicht: (2024)
Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
von: Li, Chenxu, et al.
Veröffentlicht: (2025)
von: Li, Chenxu, et al.
Veröffentlicht: (2025)
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
von: Zhu, Hengchuan, et al.
Veröffentlicht: (2025)
von: Zhu, Hengchuan, et al.
Veröffentlicht: (2025)
Proverbs Run in Pairs: Evaluating Proverb Translation Capability of Large Language Model
von: Wang, Minghan, et al.
Veröffentlicht: (2025)
von: Wang, Minghan, et al.
Veröffentlicht: (2025)
Dental-TriageBench: Benchmarking Multimodal Reasoning for Hierarchical Dental Triage
von: He, Ziyi, et al.
Veröffentlicht: (2026)
von: He, Ziyi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
von: Ge, Wentao, et al.
Veröffentlicht: (2023) -
PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain
von: Chen, Liang, et al.
Veröffentlicht: (2024) -
HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models
von: Que, Haoran, et al.
Veröffentlicht: (2024) -
MLLM-CL: Continual Learning for Multimodal Large Language Models
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025) -
RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues
von: Kuo, Tzu-Lin, et al.
Veröffentlicht: (2024)