Evaluating an evidence-guided reinforcement learning framework in aligning light-parameter large language models with decision-making cognition in psychiatric clinical reasoning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lin, Xinxin, Dai, Guangxin, Zhong, Yi, Li, Xiang, Xiao, Xue, Zhang, Yixin, Wu, Zhengdong, Zheng, Yongbo, Zhu, Runchuan, Zhao, Ming, Yu, Huizi, Wu, Shuo, Zhao, Jun, Hu, Lingming, Wang, Yumei, Yin, Ping, Chan, Joey W. Y., Chan, Ngan Yin, Chen, Sijing, Wing, Yun Kwok, Lu, Lin, Ma, Xin, Fan, Lizhou
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914310684409856
author Lin, Xinxin
Dai, Guangxin
Zhong, Yi
Li, Xiang
Xiao, Xue
Zhang, Yixin
Wu, Zhengdong
Zheng, Yongbo
Zhu, Runchuan
Zhao, Ming
Yu, Huizi
Wu, Shuo
Zhao, Jun
Hu, Lingming
Wang, Yumei
Yin, Ping
Chan, Joey W. Y.
Chan, Ngan Yin
Chen, Sijing
Wing, Yun Kwok
Lu, Lin
Ma, Xin
Fan, Lizhou
author_facet Lin, Xinxin
Dai, Guangxin
Zhong, Yi
Li, Xiang
Xiao, Xue
Zhang, Yixin
Wu, Zhengdong
Zheng, Yongbo
Zhu, Runchuan
Zhao, Ming
Yu, Huizi
Wu, Shuo
Zhao, Jun
Hu, Lingming
Wang, Yumei
Yin, Ping
Chan, Joey W. Y.
Chan, Ngan Yin
Chen, Sijing
Wing, Yun Kwok
Lu, Lin
Ma, Xin
Fan, Lizhou
contents Large language models (LLMs) hold transformative potential for medical decision support yet their application in psychiatry remains constrained by hallucinations and superficial reasoning. This limitation is particularly acute in light-parameter LLMs which are essential for privacy-preserving and efficient clinical deployment. Existing training paradigms prioritize linguistic fluency over structured clinical logic and result in a fundamental misalignment with professional diagnostic cognition. Here we introduce ClinMPO, a reinforcement learning framework designed to align the internal reasoning of LLMs with professional psychiatric practice. The framework employs a specialized reward model trained independently on a dataset derived from 4,474 psychiatry journal articles and structured according to evidence-based medicine principles. We evaluated ClinMPO on a unseen subset of the benchmark designed to isolate reasoning capabilities from rote memorization. This test set comprises items where leading large-parameter LLMs consistently fail. We compared the ClinMPO-aligned light LLM performance against a cohort of 300 medical students. The ClinMPO-tuned Qwen3-8B model achieved a diagnostic accuracy of 31.4% and surpassed the human benchmark of 30.8% on these complex cases. These results demonstrate that medical evidence-guided optimization enables light-parameter LLMs to master complex reasoning tasks. Our findings suggest that explicit cognitive alignment offers a scalable pathway to reliable and safe psychiatric decision support.
format Preprint
id arxiv_https___arxiv_org_abs_2602_06449
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Evaluating an evidence-guided reinforcement learning framework in aligning light-parameter large language models with decision-making cognition in psychiatric clinical reasoning
Lin, Xinxin
Dai, Guangxin
Zhong, Yi
Li, Xiang
Xiao, Xue
Zhang, Yixin
Wu, Zhengdong
Zheng, Yongbo
Zhu, Runchuan
Zhao, Ming
Yu, Huizi
Wu, Shuo
Zhao, Jun
Hu, Lingming
Wang, Yumei
Yin, Ping
Chan, Joey W. Y.
Chan, Ngan Yin
Chen, Sijing
Wing, Yun Kwok
Lu, Lin
Ma, Xin
Fan, Lizhou
Computation and Language
I.2.7
Large language models (LLMs) hold transformative potential for medical decision support yet their application in psychiatry remains constrained by hallucinations and superficial reasoning. This limitation is particularly acute in light-parameter LLMs which are essential for privacy-preserving and efficient clinical deployment. Existing training paradigms prioritize linguistic fluency over structured clinical logic and result in a fundamental misalignment with professional diagnostic cognition. Here we introduce ClinMPO, a reinforcement learning framework designed to align the internal reasoning of LLMs with professional psychiatric practice. The framework employs a specialized reward model trained independently on a dataset derived from 4,474 psychiatry journal articles and structured according to evidence-based medicine principles. We evaluated ClinMPO on a unseen subset of the benchmark designed to isolate reasoning capabilities from rote memorization. This test set comprises items where leading large-parameter LLMs consistently fail. We compared the ClinMPO-aligned light LLM performance against a cohort of 300 medical students. The ClinMPO-tuned Qwen3-8B model achieved a diagnostic accuracy of 31.4% and surpassed the human benchmark of 30.8% on these complex cases. These results demonstrate that medical evidence-guided optimization enables light-parameter LLMs to master complex reasoning tasks. Our findings suggest that explicit cognitive alignment offers a scalable pathway to reliable and safe psychiatric decision support.
title Evaluating an evidence-guided reinforcement learning framework in aligning light-parameter large language models with decision-making cognition in psychiatric clinical reasoning
topic Computation and Language
I.2.7
url https://arxiv.org/abs/2602.06449