Overview of the ClinicalSkillQA 2026 Shared Task on Continuous Perception and Procedural Reasoning in Clinical Skill Assessment
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Xiyang, Wei, Renxiong, Xu, Yihuai, Chen, Zhiyuan, Wu, Keying, Xiang, Jiayi, Tang, Buzhou, Ye, Yanqing, Chen, Jinyu, Zeng, Cheng, Peng, Min, Xie, Qianqian, Ananiadou, Sophia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos
di: Huang, Xiyang, et al.
Pubblicazione: (2026)
di: Huang, Xiyang, et al.
Pubblicazione: (2026)
Overview of the PromptCBLUE Shared Task in CHIP2023
di: Zhu, Wei, et al.
Pubblicazione: (2023)
di: Zhu, Wei, et al.
Pubblicazione: (2023)
ProSkill: Segment-Level Skill Assessment in Procedural Videos
di: Mazzamuto, Michele, et al.
Pubblicazione: (2026)
di: Mazzamuto, Michele, et al.
Pubblicazione: (2026)
Factual consistency evaluation of summarization in the Era of large language models
di: Luo, Zheheng, et al.
Pubblicazione: (2024)
di: Luo, Zheheng, et al.
Pubblicazione: (2024)
The Lay Person's Guide to Biomedicine: Orchestrating Large Language Models
di: Luo, Zheheng, et al.
Pubblicazione: (2024)
di: Luo, Zheheng, et al.
Pubblicazione: (2024)
CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents
di: Ossowski, Timothy, et al.
Pubblicazione: (2026)
di: Ossowski, Timothy, et al.
Pubblicazione: (2026)
MedQA-CS: Objective Structured Clinical Examination (OSCE)-Style Benchmark for Evaluating LLM Clinical Skills
di: Yao, Zonghai, et al.
Pubblicazione: (2024)
di: Yao, Zonghai, et al.
Pubblicazione: (2024)
LongDocFACTScore: Evaluating the Factuality of Long Document Abstractive Summarisation
di: Bishop, Jennifer A, et al.
Pubblicazione: (2023)
di: Bishop, Jennifer A, et al.
Pubblicazione: (2023)
Overview of the First Shared Task on Clinical Text Generation: RRG24 and "Discharge Me!"
di: Xu, Justin, et al.
Pubblicazione: (2024)
di: Xu, Justin, et al.
Pubblicazione: (2024)
Journal of Clinical Care and Skills
Pubblicazione: (2025)
Pubblicazione: (2025)
QIAS 2025: Overview of the Shared Task on Islamic Inheritance Reasoning and Knowledge Assessment
di: BOUCHEKIF, Abdessalam, et al.
Pubblicazione: (2026)
di: BOUCHEKIF, Abdessalam, et al.
Pubblicazione: (2026)
Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA
di: Chen, Guanhua, et al.
Pubblicazione: (2026)
di: Chen, Guanhua, et al.
Pubblicazione: (2026)
Sharing Skills.
di: Snoddon, Ruth V., et al.
Pubblicazione: (1987)
di: Snoddon, Ruth V., et al.
Pubblicazione: (1987)
Sharing Skills.
di: Mealy, Virginia, et al.
Pubblicazione: (1987)
di: Mealy, Virginia, et al.
Pubblicazione: (1987)
Sharing Skills.
di: Thompson, Verna, et al.
Pubblicazione: (1994)
di: Thompson, Verna, et al.
Pubblicazione: (1994)
From Scores to Skills: A Cognitive Diagnosis Framework for Evaluating Financial Large Language Models
di: Kuang, Ziyan, et al.
Pubblicazione: (2025)
di: Kuang, Ziyan, et al.
Pubblicazione: (2025)
Radon measure solutions of spherically symmetric isentropic compressible Euler equations
di: Renxiong Zhao
Pubblicazione: (2024)
di: Renxiong Zhao
Pubblicazione: (2024)
Are Large Language Models True Healthcare Jacks-of-All-Trades? Benchmarking Across Health Professions Beyond Physician Exams
di: Luo, Zheheng, et al.
Pubblicazione: (2024)
di: Luo, Zheheng, et al.
Pubblicazione: (2024)
Assessment of Rheumatology Fellows’ Skills as Clinical Teachers Through Self‐Assessment and Direct Observation
di: David Leverenz, et al.
Pubblicazione: (2025)
di: David Leverenz, et al.
Pubblicazione: (2025)
Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language Models
di: Yu, Zeping, et al.
Pubblicazione: (2025)
di: Yu, Zeping, et al.
Pubblicazione: (2025)
THCM-CAL: Temporal-Hierarchical Causal Modelling with Conformal Calibration for Clinical Risk Prediction
di: Zhang, Xin, et al.
Pubblicazione: (2025)
di: Zhang, Xin, et al.
Pubblicazione: (2025)
Student Dietitians' Perceptions and Experiences of Objective Structured Clinical Examination to Assess Communication Skills
di: Annemarie Knight, et al.
Pubblicazione: (2025)
di: Annemarie Knight, et al.
Pubblicazione: (2025)
Clinical Value of Early Motor Skill Assessment and Intervention in Autism Spectrum Disorder
di: Xue Yaqi, et al.
Pubblicazione: (2026)
di: Xue Yaqi, et al.
Pubblicazione: (2026)
Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills
di: Liu, Dawei, et al.
Pubblicazione: (2026)
di: Liu, Dawei, et al.
Pubblicazione: (2026)
How to Codesign Infographics to Teach Clinical Skills
di: Vanessa Otti, et al.
Pubblicazione: (2025)
di: Vanessa Otti, et al.
Pubblicazione: (2025)
Medical Kitchen: Transdisciplinary Clinical Skills Training
di: Jakub L. Radzikowski, et al.
Pubblicazione: (2025)
di: Jakub L. Radzikowski, et al.
Pubblicazione: (2025)
Enhancing Quantitative Reasoning Skills of Large Language Models through Dimension Perception
di: Huang, Yuncheng, et al.
Pubblicazione: (2023)
di: Huang, Yuncheng, et al.
Pubblicazione: (2023)
Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks
di: Wu, Xiyang, et al.
Pubblicazione: (2026)
di: Wu, Xiyang, et al.
Pubblicazione: (2026)
RAAR: Retrieval Augmented Agentic Reasoning for Cross-Domain Misinformation Detection
di: Liu, Zhiwei, et al.
Pubblicazione: (2026)
di: Liu, Zhiwei, et al.
Pubblicazione: (2026)
SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills
di: Lei, Yingtie, et al.
Pubblicazione: (2026)
di: Lei, Yingtie, et al.
Pubblicazione: (2026)
Interpreting Arithmetic Mechanism in Large Language Models through Comparative Neuron Analysis
di: Yu, Zeping, et al.
Pubblicazione: (2024)
di: Yu, Zeping, et al.
Pubblicazione: (2024)
How do Large Language Models Learn In-Context? Query and Key Matrices of In-Context Heads are Two Towers for Metric Learning
di: Yu, Zeping, et al.
Pubblicazione: (2024)
di: Yu, Zeping, et al.
Pubblicazione: (2024)
Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question Answering
di: Yu, Zeping, et al.
Pubblicazione: (2024)
di: Yu, Zeping, et al.
Pubblicazione: (2024)
Neuron-Level Knowledge Attribution in Large Language Models
di: Yu, Zeping, et al.
Pubblicazione: (2023)
di: Yu, Zeping, et al.
Pubblicazione: (2023)
Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing
di: Yu, Zeping, et al.
Pubblicazione: (2025)
di: Yu, Zeping, et al.
Pubblicazione: (2025)
Locate-then-Merge: Neuron-Level Parameter Fusion for Mitigating Catastrophic Forgetting in Multimodal LLMs
di: Yu, Zeping, et al.
Pubblicazione: (2025)
di: Yu, Zeping, et al.
Pubblicazione: (2025)
NOMAD: A Multi-Agent LLM System for UML Class Diagram Generation from Natural Language Requirements
di: Giannouris, Polydoros, et al.
Pubblicazione: (2025)
di: Giannouris, Polydoros, et al.
Pubblicazione: (2025)
Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2025)
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2025)
Next Token Perception Score: Analytical Assessment of your LLM Perception Skills
di: Cheng, Yu-Ang, et al.
Pubblicazione: (2025)
di: Cheng, Yu-Ang, et al.
Pubblicazione: (2025)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
di: Kabir, Mohsinul, et al.
Pubblicazione: (2026)
di: Kabir, Mohsinul, et al.
Pubblicazione: (2026)
Documenti analoghi
-
SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos
di: Huang, Xiyang, et al.
Pubblicazione: (2026) -
Overview of the PromptCBLUE Shared Task in CHIP2023
di: Zhu, Wei, et al.
Pubblicazione: (2023) -
ProSkill: Segment-Level Skill Assessment in Procedural Videos
di: Mazzamuto, Michele, et al.
Pubblicazione: (2026) -
Factual consistency evaluation of summarization in the Era of large language models
di: Luo, Zheheng, et al.
Pubblicazione: (2024) -
The Lay Person's Guide to Biomedicine: Orchestrating Large Language Models
di: Luo, Zheheng, et al.
Pubblicazione: (2024)