| _version_ | 1866901140970405888 |
|---|---|
| author | D'Urso, Enrico |
| author_facet | D'Urso, Enrico |
| contents | <p>Associated pre-print paper version: https://www.techrxiv.org/doi/full/10.36227/techrxiv.176800997.72656064/v1<br><br>This dataset contains raw inference outputs from five large language models (LLMs) on ten medical question-answering benchmarks (N=21,324 samples). Each sample includes: <br> <br> - Full chain-of-thought reasoning traces <br> - Model predictions and ground-truth labels <br> - Verbalized confidence scores (0-100) <br> - Token-level log probabilities with top-20 alternatives for each token <br> <br> Models: GPT-oss-120B (medium/high reasoning), DeepSeek-R1-Distill-32B, Qwen3-32B, Olmo-3-32B-Think <br> <br> Datasets: MedQA, MedMCQA, PubMedQA, MMLU, MMLU-Pro, MedBullets, MedExQA, AfriMedQA, MedXpertQA-R, MedXpertQA-U <br> <br> The dataset enables research on uncertainty quantification, error detection, and reasoning analysis in medical AI systems. It accompanies the paper "Hidden-uncertainty Assessment via Non-verbalized Signatures for Medical QA" (D'Urso, 2026).<br> <br> Additionally includes a robustness subset with 5 independent inference runs on 300 samples per dataset for self-consistency analysis. </p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_18130156 |
| institution | Zenodo |
| language | eng |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Hidden-uncertainty Assessment via Non-verbalized Signatures for Medical QA - Dataset D'Urso, Enrico LLM Medical Question Answering Uncertainty Quantification Trust-Abstain Systems Clinical NLP <p>Associated pre-print paper version: https://www.techrxiv.org/doi/full/10.36227/techrxiv.176800997.72656064/v1<br><br>This dataset contains raw inference outputs from five large language models (LLMs) on ten medical question-answering benchmarks (N=21,324 samples). Each sample includes: <br> <br> - Full chain-of-thought reasoning traces <br> - Model predictions and ground-truth labels <br> - Verbalized confidence scores (0-100) <br> - Token-level log probabilities with top-20 alternatives for each token <br> <br> Models: GPT-oss-120B (medium/high reasoning), DeepSeek-R1-Distill-32B, Qwen3-32B, Olmo-3-32B-Think <br> <br> Datasets: MedQA, MedMCQA, PubMedQA, MMLU, MMLU-Pro, MedBullets, MedExQA, AfriMedQA, MedXpertQA-R, MedXpertQA-U <br> <br> The dataset enables research on uncertainty quantification, error detection, and reasoning analysis in medical AI systems. It accompanies the paper "Hidden-uncertainty Assessment via Non-verbalized Signatures for Medical QA" (D'Urso, 2026).<br> <br> Additionally includes a robustness subset with 5 independent inference runs on 300 samples per dataset for self-consistency analysis. </p> |
| title | Hidden-uncertainty Assessment via Non-verbalized Signatures for Medical QA - Dataset |
| topic | LLM Medical Question Answering Uncertainty Quantification Trust-Abstain Systems Clinical NLP |
| url | https://doi.org/10.5281/zenodo.18130156 |