Saved in:
Bibliographic Details
Main Authors: Lorge, Isabelle, Joyce, Dan W., Kormilitzin, Andrey
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2404.16461
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914771190677504
author Lorge, Isabelle
Joyce, Dan W.
Kormilitzin, Andrey
author_facet Lorge, Isabelle
Joyce, Dan W.
Kormilitzin, Andrey
contents Mental health in children and adolescents has been steadily deteriorating over the past few years. The recent advent of Large Language Models (LLMs) offers much hope for cost and time efficient scaling of monitoring and intervention, yet despite specifically prevalent issues such as school bullying and eating disorders, previous studies on have not investigated performance in this domain or for open information extraction where the set of answers is not predetermined. We create a new dataset of Reddit posts from adolescents aged 12-19 annotated by expert psychiatrists for the following categories: TRAUMA, PRECARITY, CONDITION, SYMPTOMS, SUICIDALITY and TREATMENT and compare expert labels to annotations from two top performing LLMs (GPT3.5 and GPT4). In addition, we create two synthetic datasets to assess whether LLMs perform better when annotating data as they generate it. We find GPT4 to be on par with human inter-annotator agreement and performance on synthetic data to be substantially higher, however we find the model still occasionally errs on issues of negation and factuality and higher performance on synthetic data is driven by greater complexity of real data rather than inherent advantage.
format Preprint
id arxiv_https___arxiv_org_abs_2404_16461
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Large Language Models Perform on Par with Experts Identifying Mental Health Factors in Adolescent Online Forums
Lorge, Isabelle
Joyce, Dan W.
Kormilitzin, Andrey
Computation and Language
Mental health in children and adolescents has been steadily deteriorating over the past few years. The recent advent of Large Language Models (LLMs) offers much hope for cost and time efficient scaling of monitoring and intervention, yet despite specifically prevalent issues such as school bullying and eating disorders, previous studies on have not investigated performance in this domain or for open information extraction where the set of answers is not predetermined. We create a new dataset of Reddit posts from adolescents aged 12-19 annotated by expert psychiatrists for the following categories: TRAUMA, PRECARITY, CONDITION, SYMPTOMS, SUICIDALITY and TREATMENT and compare expert labels to annotations from two top performing LLMs (GPT3.5 and GPT4). In addition, we create two synthetic datasets to assess whether LLMs perform better when annotating data as they generate it. We find GPT4 to be on par with human inter-annotator agreement and performance on synthetic data to be substantially higher, however we find the model still occasionally errs on issues of negation and factuality and higher performance on synthetic data is driven by greater complexity of real data rather than inherent advantage.
title Large Language Models Perform on Par with Experts Identifying Mental Health Factors in Adolescent Online Forums
topic Computation and Language
url https://arxiv.org/abs/2404.16461