Detecting PTSD in Clinical Interviews: A Comparative Analysis of NLP Methods and Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Feng, Ben-Zeev, Dror, Sparks, Gillian, Kadakia, Arya, Cohen, Trevor
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917187380314112
author Chen, Feng
Ben-Zeev, Dror
Sparks, Gillian
Kadakia, Arya
Cohen, Trevor
author_facet Chen, Feng
Ben-Zeev, Dror
Sparks, Gillian
Kadakia, Arya
Cohen, Trevor
contents Post-Traumatic Stress Disorder (PTSD) remains underdiagnosed in clinical settings, presenting opportunities for automated detection to identify patients. This study evaluates natural language processing approaches for detecting PTSD from clinical interview transcripts. We compared general and mental health-specific transformer models (BERT/RoBERTa), embedding-based methods (SentenceBERT/LLaMA), and large language model prompting strategies (zero-shot/few-shot/chain-of-thought) using the DAIC-WOZ dataset. Domain-specific end-to-end models significantly outperformed general models (Mental-RoBERTa AUPRC=0.675+/-0.084 vs. RoBERTa-base 0.599+/-0.145). SentenceBERT embeddings with neural networks achieved the highest overall performance (AUPRC=0.758+/-0.128). Few-shot prompting using DSM-5 criteria yielded competitive results with two examples (AUPRC=0.737). Performance varied significantly across symptom severity and comorbidity status with depression, with higher accuracy for severe PTSD cases and patients with comorbid depression. Our findings highlight the potential of domain-adapted embeddings and LLMs for scalable screening while underscoring the need for improved detection of nuanced presentations and offering insights for developing clinically viable AI tools for PTSD assessment.
format Preprint
id arxiv_https___arxiv_org_abs_2504_01216
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Detecting PTSD in Clinical Interviews: A Comparative Analysis of NLP Methods and Large Language Models
Chen, Feng
Ben-Zeev, Dror
Sparks, Gillian
Kadakia, Arya
Cohen, Trevor
Computation and Language
Artificial Intelligence
Machine Learning
Post-Traumatic Stress Disorder (PTSD) remains underdiagnosed in clinical settings, presenting opportunities for automated detection to identify patients. This study evaluates natural language processing approaches for detecting PTSD from clinical interview transcripts. We compared general and mental health-specific transformer models (BERT/RoBERTa), embedding-based methods (SentenceBERT/LLaMA), and large language model prompting strategies (zero-shot/few-shot/chain-of-thought) using the DAIC-WOZ dataset. Domain-specific end-to-end models significantly outperformed general models (Mental-RoBERTa AUPRC=0.675+/-0.084 vs. RoBERTa-base 0.599+/-0.145). SentenceBERT embeddings with neural networks achieved the highest overall performance (AUPRC=0.758+/-0.128). Few-shot prompting using DSM-5 criteria yielded competitive results with two examples (AUPRC=0.737). Performance varied significantly across symptom severity and comorbidity status with depression, with higher accuracy for severe PTSD cases and patients with comorbid depression. Our findings highlight the potential of domain-adapted embeddings and LLMs for scalable screening while underscoring the need for improved detection of nuanced presentations and offering insights for developing clinically viable AI tools for PTSD assessment.
title Detecting PTSD in Clinical Interviews: A Comparative Analysis of NLP Methods and Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2504.01216