Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process Rewards

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yun, Jaehoon, Sohn, Jiwoong, Park, Jungwoo, Kim, Hyunjae, Tang, Xiangru, Shao, Yanjun, Koo, Yonghoe, Ko, Minhyeok, Chen, Qingyu, Gerstein, Mark, Moor, Michael, Kang, Jaewoo
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918145325793280
author Yun, Jaehoon
Sohn, Jiwoong
Park, Jungwoo
Kim, Hyunjae
Tang, Xiangru
Shao, Yanjun
Koo, Yonghoe
Ko, Minhyeok
Chen, Qingyu
Gerstein, Mark
Moor, Michael
Kang, Jaewoo
author_facet Yun, Jaehoon
Sohn, Jiwoong
Park, Jungwoo
Kim, Hyunjae
Tang, Xiangru
Shao, Yanjun
Koo, Yonghoe
Ko, Minhyeok
Chen, Qingyu
Gerstein, Mark
Moor, Michael
Kang, Jaewoo
contents Large language models have shown promise in clinical decision making, but current approaches struggle to localize and correct errors at specific steps of the reasoning process. This limitation is critical in medicine, where identifying and addressing reasoning errors is essential for accurate diagnosis and effective patient care. We introduce Med-PRM, a process reward modeling framework that leverages retrieval-augmented generation to verify each reasoning step against established medical knowledge bases. By verifying intermediate reasoning steps with evidence retrieved from clinical guidelines and literature, our model can precisely assess the reasoning quality in a fine-grained manner. Evaluations on five medical QA benchmarks and two open-ended diagnostic tasks demonstrate that Med-PRM achieves state-of-the-art performance, with improving the performance of base models by up to 13.50% using Med-PRM. Moreover, we demonstrate the generality of Med-PRM by integrating it in a plug-and-play fashion with strong policy models such as Meerkat, achieving over 80\% accuracy on MedQA for the first time using small-scale models of 8 billion parameters. Our code and data are available at: https://med-prm.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2506_11474
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process Rewards
Yun, Jaehoon
Sohn, Jiwoong
Park, Jungwoo
Kim, Hyunjae
Tang, Xiangru
Shao, Yanjun
Koo, Yonghoe
Ko, Minhyeok
Chen, Qingyu
Gerstein, Mark
Moor, Michael
Kang, Jaewoo
Computation and Language
Large language models have shown promise in clinical decision making, but current approaches struggle to localize and correct errors at specific steps of the reasoning process. This limitation is critical in medicine, where identifying and addressing reasoning errors is essential for accurate diagnosis and effective patient care. We introduce Med-PRM, a process reward modeling framework that leverages retrieval-augmented generation to verify each reasoning step against established medical knowledge bases. By verifying intermediate reasoning steps with evidence retrieved from clinical guidelines and literature, our model can precisely assess the reasoning quality in a fine-grained manner. Evaluations on five medical QA benchmarks and two open-ended diagnostic tasks demonstrate that Med-PRM achieves state-of-the-art performance, with improving the performance of base models by up to 13.50% using Med-PRM. Moreover, we demonstrate the generality of Med-PRM by integrating it in a plug-and-play fashion with strong policy models such as Meerkat, achieving over 80\% accuracy on MedQA for the first time using small-scale models of 8 billion parameters. Our code and data are available at: https://med-prm.github.io/
title Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process Rewards
topic Computation and Language
url https://arxiv.org/abs/2506.11474