Saved in:
Bibliographic Details
Main Authors: Kaliosis, Panagiotis, Ganesan, Adithya V, Kjell, Oscar N. E., Ringwald, Whitney, Feltman, Scott, Carr, Melissa A., Samaras, Dimitris, Ruggero, Camilo, Luft, Benjamin J., Kotov, Roman, Schwartz, Andrew H.
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.06015
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912882012192768
author Kaliosis, Panagiotis
Ganesan, Adithya V
Kjell, Oscar N. E.
Ringwald, Whitney
Feltman, Scott
Carr, Melissa A.
Samaras, Dimitris
Ruggero, Camilo
Luft, Benjamin J.
Kotov, Roman
Schwartz, Andrew H.
author_facet Kaliosis, Panagiotis
Ganesan, Adithya V
Kjell, Oscar N. E.
Ringwald, Whitney
Feltman, Scott
Carr, Melissa A.
Samaras, Dimitris
Ruggero, Camilo
Luft, Benjamin J.
Kotov, Roman
Schwartz, Andrew H.
contents Large language models (LLMs) are increasingly being used in a zero-shot fashion to assess mental health conditions, yet we have limited knowledge on what factors affect their accuracy. In this study, we utilize a clinical dataset of natural language narratives and self-reported PTSD severity scores from 1,437 individuals to comprehensively evaluate the performance of 11 state-of-the-art LLMs. To understand the factors affecting accuracy, we systematically varied (i) contextual knowledge like subscale definitions, distribution summary, and interview questions, and (ii) modeling strategies including zero-shot vs few shot, amount of reasoning effort, model sizes, structured subscales vs direct scalar prediction, output rescaling and nine ensemble methods. Our findings indicate that (a) LLMs are most accurate when provided with detailed construct definitions and context of the narrative; (b) increased reasoning effort leads to better estimation accuracy; (c) performance of open-weight models (Llama, Deepseek), plateau beyond 70B parameters while closed-weight (o3-mini, gpt-5) models improve with newer generations; and (d) best performance is achieved when ensembling a supervised model with the zero-shot LLMs. Taken together, the results suggest choice of contextual knowledge and modeling strategies is important for deploying LLMs to accurately assess mental health.
format Preprint
id arxiv_https___arxiv_org_abs_2602_06015
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Systematic Evaluation of Large Language Models for PTSD Severity Estimation: The Role of Contextual Knowledge and Modeling Strategies
Kaliosis, Panagiotis
Ganesan, Adithya V
Kjell, Oscar N. E.
Ringwald, Whitney
Feltman, Scott
Carr, Melissa A.
Samaras, Dimitris
Ruggero, Camilo
Luft, Benjamin J.
Kotov, Roman
Schwartz, Andrew H.
Computation and Language
Large language models (LLMs) are increasingly being used in a zero-shot fashion to assess mental health conditions, yet we have limited knowledge on what factors affect their accuracy. In this study, we utilize a clinical dataset of natural language narratives and self-reported PTSD severity scores from 1,437 individuals to comprehensively evaluate the performance of 11 state-of-the-art LLMs. To understand the factors affecting accuracy, we systematically varied (i) contextual knowledge like subscale definitions, distribution summary, and interview questions, and (ii) modeling strategies including zero-shot vs few shot, amount of reasoning effort, model sizes, structured subscales vs direct scalar prediction, output rescaling and nine ensemble methods. Our findings indicate that (a) LLMs are most accurate when provided with detailed construct definitions and context of the narrative; (b) increased reasoning effort leads to better estimation accuracy; (c) performance of open-weight models (Llama, Deepseek), plateau beyond 70B parameters while closed-weight (o3-mini, gpt-5) models improve with newer generations; and (d) best performance is achieved when ensembling a supervised model with the zero-shot LLMs. Taken together, the results suggest choice of contextual knowledge and modeling strategies is important for deploying LLMs to accurately assess mental health.
title A Systematic Evaluation of Large Language Models for PTSD Severity Estimation: The Role of Contextual Knowledge and Modeling Strategies
topic Computation and Language
url https://arxiv.org/abs/2602.06015