MentraSuite: Post-Training Large Language Models for Mental Health Reasoning and Assessment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Mengxi, Yang, Kailai, Zhao, Pengde, Zhang, Enze, Kuang, Ziyan, Liu, Zhiwei, Han, Weiguang, Liao, Shu, Huang, Lianting, Hu, Jinpeng, Peng, Min, Xie, Qianqian, Ananiadou, Sophia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918249496576000
author Xiao, Mengxi
Yang, Kailai
Zhao, Pengde
Zhang, Enze
Kuang, Ziyan
Liu, Zhiwei
Han, Weiguang
Liao, Shu
Huang, Lianting
Hu, Jinpeng
Peng, Min
Xie, Qianqian
Ananiadou, Sophia
author_facet Xiao, Mengxi
Yang, Kailai
Zhao, Pengde
Zhang, Enze
Kuang, Ziyan
Liu, Zhiwei
Han, Weiguang
Liao, Shu
Huang, Lianting
Hu, Jinpeng
Peng, Min
Xie, Qianqian
Ananiadou, Sophia
contents Mental health disorders affect hundreds of millions globally, and the Web now serves as a primary medium for accessing support, information, and assessment. Large language models (LLMs) offer scalable and accessible assistance, yet their deployment in mental-health settings remains risky when their reasoning is incomplete, inconsistent, or ungrounded. Existing psychological LLMs emphasize emotional understanding or knowledge recall but overlook the step-wise, clinically aligned reasoning required for appraisal, diagnosis, intervention planning, abstraction, and verification. To address these issues, we introduce MentraSuite, a unified framework for advancing reliable mental-health reasoning. We propose MentraBench, a comprehensive benchmark spanning five core reasoning aspects, six tasks, and 13 datasets, evaluating both task performance and reasoning quality across five dimensions: conciseness, coherence, hallucination avoidance, task understanding, and internal consistency. We further present Mindora, a post-trained model optimized through a hybrid SFT-RL framework with an inconsistency-detection reward to enforce faithful and coherent reasoning. To support training, we construct high-quality trajectories using a novel reasoning trajectory generation strategy, that strategically filters difficult samples and applies a structured, consistency-oriented rewriting process to produce concise, readable, and well-balanced trajectories. Across 20 evaluated LLMs, Mindora achieves the highest average performance on MentraBench and shows remarkable performances in reasoning reliability, demonstrating its effectiveness for complex mental-health scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2512_09636
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MentraSuite: Post-Training Large Language Models for Mental Health Reasoning and Assessment
Xiao, Mengxi
Yang, Kailai
Zhao, Pengde
Zhang, Enze
Kuang, Ziyan
Liu, Zhiwei
Han, Weiguang
Liao, Shu
Huang, Lianting
Hu, Jinpeng
Peng, Min
Xie, Qianqian
Ananiadou, Sophia
Computation and Language
Mental health disorders affect hundreds of millions globally, and the Web now serves as a primary medium for accessing support, information, and assessment. Large language models (LLMs) offer scalable and accessible assistance, yet their deployment in mental-health settings remains risky when their reasoning is incomplete, inconsistent, or ungrounded. Existing psychological LLMs emphasize emotional understanding or knowledge recall but overlook the step-wise, clinically aligned reasoning required for appraisal, diagnosis, intervention planning, abstraction, and verification. To address these issues, we introduce MentraSuite, a unified framework for advancing reliable mental-health reasoning. We propose MentraBench, a comprehensive benchmark spanning five core reasoning aspects, six tasks, and 13 datasets, evaluating both task performance and reasoning quality across five dimensions: conciseness, coherence, hallucination avoidance, task understanding, and internal consistency. We further present Mindora, a post-trained model optimized through a hybrid SFT-RL framework with an inconsistency-detection reward to enforce faithful and coherent reasoning. To support training, we construct high-quality trajectories using a novel reasoning trajectory generation strategy, that strategically filters difficult samples and applies a structured, consistency-oriented rewriting process to produce concise, readable, and well-balanced trajectories. Across 20 evaluated LLMs, Mindora achieves the highest average performance on MentraBench and shows remarkable performances in reasoning reliability, demonstrating its effectiveness for complex mental-health scenarios.
title MentraSuite: Post-Training Large Language Models for Mental Health Reasoning and Assessment
topic Computation and Language
url https://arxiv.org/abs/2512.09636