AnesSuite: A Comprehensive Benchmark and Dataset Suite for Anesthesiology Reasoning in LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Feng, Xiang, Jiang, Wentao, Wang, Zengmao, Luo, Yong, Xu, Pingbo, Yu, Baosheng, Jin, Hua, Zhang, Jing |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
REX-RAG: Reasoning Exploration with Policy Correction in Retrieval-Augmented Generation
par: Jiang, Wentao, et autres
Publié: (2025)
par: Jiang, Wentao, et autres
Publié: (2025)
ConceptPsy:A Benchmark Suite with Conceptual Comprehensiveness in Psychology
par: Zhang, Junlei, et autres
Publié: (2023)
par: Zhang, Junlei, et autres
Publié: (2023)
Benchmarking Hindi LLMs: A New Suite of Datasets and a Comparative Analysis
par: Kamath, Anusha, et autres
Publié: (2025)
par: Kamath, Anusha, et autres
Publié: (2025)
Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks
par: Warner, Benjamin, et autres
Publié: (2026)
par: Warner, Benjamin, et autres
Publié: (2026)
A Benchmark Suite of Reddit-Derived Datasets for Mental Health Detection
par: Hasan, Khalid, et autres
Publié: (2026)
par: Hasan, Khalid, et autres
Publié: (2026)
LogEval: A Comprehensive Benchmark Suite for Large Language Models In Log Analysis
par: Cui, Tianyu, et autres
Publié: (2024)
par: Cui, Tianyu, et autres
Publié: (2024)
GenderBench: Evaluation Suite for Gender Biases in LLMs
par: Pikuliak, Matúš
Publié: (2025)
par: Pikuliak, Matúš
Publié: (2025)
AgriGPT-VL: Agricultural Vision-Language Understanding Suite
par: Yang, Bo, et autres
Publié: (2025)
par: Yang, Bo, et autres
Publié: (2025)
KFinEval-Pilot: A Comprehensive Benchmark Suite for Korean Financial Language Understanding
par: Hwang, Bokwang, et autres
Publié: (2025)
par: Hwang, Bokwang, et autres
Publié: (2025)
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
par: Huang, Xu, et autres
Publié: (2025)
par: Huang, Xu, et autres
Publié: (2025)
Verifiable Natural Language to Linear Temporal Logic Translation: A Benchmark Dataset and Evaluation Suite
par: English, William H, et autres
Publié: (2025)
par: English, William H, et autres
Publié: (2025)
Med42-v2: A Suite of Clinical LLMs
par: Christophe, Clément, et autres
Publié: (2024)
par: Christophe, Clément, et autres
Publié: (2024)
Towards Training A Chinese Large Language Model for Anesthesiology
par: Wang, Zhonghai, et autres
Publié: (2024)
par: Wang, Zhonghai, et autres
Publié: (2024)
TigerCoder: A Novel Suite of LLMs for Code Generation in Bangla
par: Raihan, Nishat, et autres
Publié: (2025)
par: Raihan, Nishat, et autres
Publié: (2025)
FormFactory: An Interactive Benchmarking Suite for Multimodal Form-Filling Agents
par: Li, Bobo, et autres
Publié: (2025)
par: Li, Bobo, et autres
Publié: (2025)
MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks
par: Ruan, Junhao, et autres
Publié: (2026)
par: Ruan, Junhao, et autres
Publié: (2026)
MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation
par: Yao, Jihan, et autres
Publié: (2025)
par: Yao, Jihan, et autres
Publié: (2025)
RADIUS: Ranking, Distribution, and Significance - A Comprehensive Alignment Suite for Survey Simulation
par: Łajewska, Weronika, et autres
Publié: (2026)
par: Łajewska, Weronika, et autres
Publié: (2026)
LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems
par: Wu, Siyu, et autres
Publié: (2026)
par: Wu, Siyu, et autres
Publié: (2026)
SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages
par: Lovenia, Holy, et autres
Publié: (2024)
par: Lovenia, Holy, et autres
Publié: (2024)
LoraxBench: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languages
par: Aji, Alham Fikri, et autres
Publié: (2025)
par: Aji, Alham Fikri, et autres
Publié: (2025)
Diagnosing Translated Benchmarks: An Automated Quality Assurance Study of the EU20 Benchmark Suite
par: Thellmann, Klaudia, et autres
Publié: (2026)
par: Thellmann, Klaudia, et autres
Publié: (2026)
MentraSuite: Post-Training Large Language Models for Mental Health Reasoning and Assessment
par: Xiao, Mengxi, et autres
Publié: (2025)
par: Xiao, Mengxi, et autres
Publié: (2025)
The CodeInverter Suite: Control-Flow and Data-Mapping Augmented Binary Decompilation with LLMs
par: Liu, Peipei, et autres
Publié: (2025)
par: Liu, Peipei, et autres
Publié: (2025)
TALES: Text Adventure Learning Environment Suite
par: Cui, Christopher Zhang, et autres
Publié: (2025)
par: Cui, Christopher Zhang, et autres
Publié: (2025)
MoTime: A Dataset Suite for Multimodal Time Series Forecasting
par: Zhou, Xin, et autres
Publié: (2025)
par: Zhou, Xin, et autres
Publié: (2025)
VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
par: Huang, Ziqi, et autres
Publié: (2024)
par: Huang, Ziqi, et autres
Publié: (2024)
A 200-Line Python Micro-Benchmark Suite for NISQ Circuit Compilers
par: Merilehto, Juhani
Publié: (2025)
par: Merilehto, Juhani
Publié: (2025)
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
par: Bragg, Jonathan, et autres
Publié: (2025)
par: Bragg, Jonathan, et autres
Publié: (2025)
FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling
par: Xu, Guixian, et autres
Publié: (2026)
par: Xu, Guixian, et autres
Publié: (2026)
ViLLM-Eval: A Comprehensive Evaluation Suite for Vietnamese Large Language Models
par: Nguyen, Trong-Hieu, et autres
Publié: (2024)
par: Nguyen, Trong-Hieu, et autres
Publié: (2024)
PsyEval: A Suite of Mental Health Related Tasks for Evaluating Large Language Models
par: Jin, Haoan, et autres
Publié: (2023)
par: Jin, Haoan, et autres
Publié: (2023)
The Zamba2 Suite: Technical Report
par: Glorioso, Paolo, et autres
Publié: (2024)
par: Glorioso, Paolo, et autres
Publié: (2024)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
par: Kim, Eunsu, et autres
Publié: (2025)
par: Kim, Eunsu, et autres
Publié: (2025)
AMNESIA: A Large Scale Medical Unlearning Benchmark Suite with Disease-Informed Analysis
par: Davoudi, Saeedeh, et autres
Publié: (2026)
par: Davoudi, Saeedeh, et autres
Publié: (2026)
CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics
par: Somasekharan, Nithin, et autres
Publié: (2025)
par: Somasekharan, Nithin, et autres
Publié: (2025)
ViStoryBench: Comprehensive Benchmark Suite for Story Visualization
par: Zhuang, Cailin, et autres
Publié: (2025)
par: Zhuang, Cailin, et autres
Publié: (2025)
CFBench: A Comprehensive Constraints-Following Benchmark for LLMs
par: Zhang, Tao, et autres
Publié: (2024)
par: Zhang, Tao, et autres
Publié: (2024)
LeanCat: A Benchmark Suite for Formal Category Theory in Lean (Part I: 1-Categories)
par: Xu, Rongge, et autres
Publié: (2025)
par: Xu, Rongge, et autres
Publié: (2025)
MatNexus: A Comprehensive Text Mining and Analysis Suite for Materials Discover
par: Zhang, Lei, et autres
Publié: (2023)
par: Zhang, Lei, et autres
Publié: (2023)
Documents similaires
-
REX-RAG: Reasoning Exploration with Policy Correction in Retrieval-Augmented Generation
par: Jiang, Wentao, et autres
Publié: (2025) -
ConceptPsy:A Benchmark Suite with Conceptual Comprehensiveness in Psychology
par: Zhang, Junlei, et autres
Publié: (2023) -
Benchmarking Hindi LLMs: A New Suite of Datasets and a Comparative Analysis
par: Kamath, Anusha, et autres
Publié: (2025) -
Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks
par: Warner, Benjamin, et autres
Publié: (2026) -
A Benchmark Suite of Reddit-Derived Datasets for Mental Health Detection
par: Hasan, Khalid, et autres
Publié: (2026)