Hidden-uncertainty Assessment via Non-verbalized Signatures for Medical QA - Dataset

Fuente: Zenodo
Saved in:
Bibliographic Details
Main Author: D'Urso, Enrico
Format: Recurso digital
Language:English
Published: Zenodo 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866901140970405888
author D'Urso, Enrico
author_facet D'Urso, Enrico
contents <p>Associated pre-print paper version: https://www.techrxiv.org/doi/full/10.36227/techrxiv.176800997.72656064/v1<br><br>This dataset contains raw inference outputs from five large language models (LLMs) on ten medical question-answering benchmarks (N=21,324 samples). Each sample includes:                                                                <br>                                                                                                                                                                                                                                           <br>  - Full chain-of-thought reasoning traces                                                                                                                                                                                                 <br>  - Model predictions and ground-truth labels                                                                                                                                                                                              <br>  - Verbalized confidence scores (0-100)                                                                                                                                                                                                   <br>  - Token-level log probabilities with top-20 alternatives for each token                                                                                                                                                                  <br>                                                                                                                                                                                                                                           <br>  Models: GPT-oss-120B (medium/high reasoning), DeepSeek-R1-Distill-32B, Qwen3-32B, Olmo-3-32B-Think                                                                                                                                       <br>                                                                                                                                                                                                                                           <br>  Datasets: MedQA, MedMCQA, PubMedQA, MMLU, MMLU-Pro, MedBullets, MedExQA, AfriMedQA, MedXpertQA-R, MedXpertQA-U                                                                                                                           <br>                                                                                                                                                                                                                                           <br>  The dataset enables research on uncertainty quantification, error detection, and reasoning analysis in medical AI systems. It accompanies the paper "Hidden-uncertainty Assessment via Non-verbalized Signatures for Medical QA" (D'Urso, 2026).<br>                                                                                                                                                                                                                                           <br>  Additionally includes a robustness subset with 5 independent inference runs on 300 samples per dataset for self-consistency analysis.   </p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_18130156
institution Zenodo
language eng
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Hidden-uncertainty Assessment via Non-verbalized Signatures for Medical QA - Dataset
D'Urso, Enrico
LLM
Medical Question Answering
Uncertainty Quantification
Trust-Abstain Systems
Clinical NLP
<p>Associated pre-print paper version: https://www.techrxiv.org/doi/full/10.36227/techrxiv.176800997.72656064/v1<br><br>This dataset contains raw inference outputs from five large language models (LLMs) on ten medical question-answering benchmarks (N=21,324 samples). Each sample includes:                                                                <br>                                                                                                                                                                                                                                           <br>  - Full chain-of-thought reasoning traces                                                                                                                                                                                                 <br>  - Model predictions and ground-truth labels                                                                                                                                                                                              <br>  - Verbalized confidence scores (0-100)                                                                                                                                                                                                   <br>  - Token-level log probabilities with top-20 alternatives for each token                                                                                                                                                                  <br>                                                                                                                                                                                                                                           <br>  Models: GPT-oss-120B (medium/high reasoning), DeepSeek-R1-Distill-32B, Qwen3-32B, Olmo-3-32B-Think                                                                                                                                       <br>                                                                                                                                                                                                                                           <br>  Datasets: MedQA, MedMCQA, PubMedQA, MMLU, MMLU-Pro, MedBullets, MedExQA, AfriMedQA, MedXpertQA-R, MedXpertQA-U                                                                                                                           <br>                                                                                                                                                                                                                                           <br>  The dataset enables research on uncertainty quantification, error detection, and reasoning analysis in medical AI systems. It accompanies the paper "Hidden-uncertainty Assessment via Non-verbalized Signatures for Medical QA" (D'Urso, 2026).<br>                                                                                                                                                                                                                                           <br>  Additionally includes a robustness subset with 5 independent inference runs on 300 samples per dataset for self-consistency analysis.   </p>
title Hidden-uncertainty Assessment via Non-verbalized Signatures for Medical QA - Dataset
topic LLM
Medical Question Answering
Uncertainty Quantification
Trust-Abstain Systems
Clinical NLP
url https://doi.org/10.5281/zenodo.18130156