CMB: A Comprehensive Medical Benchmark in Chinese
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xidong, Chen, Guiming Hardy, Song, Dingjie, Zhang, Zhiyi, Chen, Zhihong, Xiao, Qingying, Jiang, Feng, Li, Jianquan, Wan, Xiang, Wang, Benyou, Li, Haizhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MileBench: Benchmarking MLLMs in Long Context
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
von: Chen, Guiming Hardy, et al.
Veröffentlicht: (2024)
von: Chen, Guiming Hardy, et al.
Veröffentlicht: (2024)
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
von: Ge, Wentao, et al.
Veröffentlicht: (2023)
von: Ge, Wentao, et al.
Veröffentlicht: (2023)
HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs
von: Chen, Junying, et al.
Veröffentlicht: (2023)
von: Chen, Junying, et al.
Veröffentlicht: (2023)
Enabling Doctor-Centric Medical AI with LLMs through Workflow-Aligned Tasks and Benchmarks
von: Xie, Wenya, et al.
Veröffentlicht: (2025)
von: Xie, Wenya, et al.
Veröffentlicht: (2025)
LLMs for Doctors: Leveraging Medical LLMs to Assist Doctors, Not Replace Them
von: Xie, Wenya, et al.
Veröffentlicht: (2024)
von: Xie, Wenya, et al.
Veröffentlicht: (2024)
Humans or LLMs as the Judge? A Study on Judgement Biases
von: Chen, Guiming Hardy, et al.
Veröffentlicht: (2024)
von: Chen, Guiming Hardy, et al.
Veröffentlicht: (2024)
Quantifying Self-diagnostic Atomic Knowledge in Chinese Medical Foundation Model: A Computational Analysis
von: Fan, Yaxin, et al.
Veröffentlicht: (2023)
von: Fan, Yaxin, et al.
Veröffentlicht: (2023)
Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
Apollo: A Lightweight Multilingual Medical LLM towards Democratizing Medical AI to 6B People
von: Wang, Xidong, et al.
Veröffentlicht: (2024)
von: Wang, Xidong, et al.
Veröffentlicht: (2024)
HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
von: Chen, Junying, et al.
Veröffentlicht: (2024)
von: Chen, Junying, et al.
Veröffentlicht: (2024)
AceGPT, Localizing Large Language Models in Arabic
von: Huang, Huang, et al.
Veröffentlicht: (2023)
von: Huang, Huang, et al.
Veröffentlicht: (2023)
ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese Medicine
von: Chen, Junying, et al.
Veröffentlicht: (2025)
von: Chen, Junying, et al.
Veröffentlicht: (2025)
CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis
von: Chen, Junying, et al.
Veröffentlicht: (2024)
von: Chen, Junying, et al.
Veröffentlicht: (2024)
LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture
von: Wang, Xidong, et al.
Veröffentlicht: (2024)
von: Wang, Xidong, et al.
Veröffentlicht: (2024)
PrinciplismQA: A Philosophy-Grounded Approach to Assessing LLM-Human Clinical Medical Ethics Alignment
von: Hong, Chang, et al.
Veröffentlicht: (2025)
von: Hong, Chang, et al.
Veröffentlicht: (2025)
Beyond Binary: Towards Fine-Grained LLM-Generated Text Detection via Role Recognition and Involvement Measurement
von: Cheng, Zihao, et al.
Veröffentlicht: (2024)
von: Cheng, Zihao, et al.
Veröffentlicht: (2024)
Roadmap towards Superhuman Speech Understanding using Large Language Models
von: Bu, Fan, et al.
Veröffentlicht: (2024)
von: Bu, Fan, et al.
Veröffentlicht: (2024)
Efficiently Democratizing Medical LLMs for 50 Languages via a Mixture of Language Family Experts
von: Zheng, Guorui, et al.
Veröffentlicht: (2024)
von: Zheng, Guorui, et al.
Veröffentlicht: (2024)
Online Training of Large Language Models: Learn while chatting
von: Liang, Juhao, et al.
Veröffentlicht: (2024)
von: Liang, Juhao, et al.
Veröffentlicht: (2024)
Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging
von: Cai, Zhenyang, et al.
Veröffentlicht: (2024)
von: Cai, Zhenyang, et al.
Veröffentlicht: (2024)
From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test
von: Dai, Xunlian, et al.
Veröffentlicht: (2025)
von: Dai, Xunlian, et al.
Veröffentlicht: (2025)
A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models
von: Jing, Liqiang, et al.
Veröffentlicht: (2025)
von: Jing, Liqiang, et al.
Veröffentlicht: (2025)
MedBench: A Comprehensive, Standardized, and Reliable Benchmarking System for Evaluating Chinese Medical Large Language Models
von: Liu, Mianxin, et al.
Veröffentlicht: (2024)
von: Liu, Mianxin, et al.
Veröffentlicht: (2024)
Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
Do LLMs Triage Like Clinicians? A Dynamic Study of Outpatient Referral
von: Liu, Xiaoxiao, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoxiao, et al.
Veröffentlicht: (2025)
HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
von: Chen, Junying, et al.
Veröffentlicht: (2024)
von: Chen, Junying, et al.
Veröffentlicht: (2024)
PlatoLM: Teaching LLMs in Multi-Round Dialogue via a User Simulator
von: Kong, Chuyi, et al.
Veröffentlicht: (2023)
von: Kong, Chuyi, et al.
Veröffentlicht: (2023)
PediaBench: A Comprehensive Chinese Pediatric Dataset for Benchmarking Large Language Models
von: Zhang, Qian, et al.
Veröffentlicht: (2024)
von: Zhang, Qian, et al.
Veröffentlicht: (2024)
Advancing Topic Segmentation and Outline Generation in Chinese Texts: The Paragraph-level Topic Representation, Corpus, and Benchmark
von: Jiang, Feng, et al.
Veröffentlicht: (2023)
von: Jiang, Feng, et al.
Veröffentlicht: (2023)
Bridging Research and Readers: A Multi-Modal Automated Academic Papers Interpretation System
von: Jiang, Feng, et al.
Veröffentlicht: (2024)
von: Jiang, Feng, et al.
Veröffentlicht: (2024)
EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models
von: Zhou, Li, et al.
Veröffentlicht: (2025)
von: Zhou, Li, et al.
Veröffentlicht: (2025)
INSEva: A Comprehensive Chinese Benchmark for Large Language Models in Insurance
von: Chen, Shisong, et al.
Veröffentlicht: (2025)
von: Chen, Shisong, et al.
Veröffentlicht: (2025)
WaveMind: Towards a Conversational EEG Foundation Model Aligned to Textual and Visual Modalities
von: Zeng, Ziyi, et al.
Veröffentlicht: (2025)
von: Zeng, Ziyi, et al.
Veröffentlicht: (2025)
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
AgriEval: A Comprehensive Chinese Agricultural Benchmark for Large Language Models
von: Yan, Lian, et al.
Veröffentlicht: (2025)
von: Yan, Lian, et al.
Veröffentlicht: (2025)
LLMs Could Autonomously Learn Without External Supervision
von: Ji, Ke, et al.
Veröffentlicht: (2024)
von: Ji, Ke, et al.
Veröffentlicht: (2024)
Soundwave: Less is More for Speech-Text Alignment in LLMs
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
Is ChatGPT Involved in Texts? Measure the Polish Ratio to Detect ChatGPT-Generated Text
von: Yang, Lingyi, et al.
Veröffentlicht: (2023)
von: Yang, Lingyi, et al.
Veröffentlicht: (2023)
Falcon: A Comprehensive Chinese Text-to-SQL Benchmark for Enterprise-Grade Evaluation
von: Luo, Wenzhen, et al.
Veröffentlicht: (2025)
von: Luo, Wenzhen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MileBench: Benchmarking MLLMs in Long Context
von: Song, Dingjie, et al.
Veröffentlicht: (2024) -
ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
von: Chen, Guiming Hardy, et al.
Veröffentlicht: (2024) -
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
von: Ge, Wentao, et al.
Veröffentlicht: (2023) -
HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs
von: Chen, Junying, et al.
Veröffentlicht: (2023) -
Enabling Doctor-Centric Medical AI with LLMs through Workflow-Aligned Tasks and Benchmarks
von: Xie, Wenya, et al.
Veröffentlicht: (2025)