Salvato in:
Dettagli Bibliografici
Autori principali: Tank, Chayan, Mehta, Shaina, Pol, Sarthak, Katoch, Vinayak, Anand, Avinash, Jaiswal, Raj, Shah, Rajiv Ratn
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:https://arxiv.org/abs/2412.01353
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910751870943232
author Tank, Chayan
Mehta, Shaina
Pol, Sarthak
Katoch, Vinayak
Anand, Avinash
Jaiswal, Raj
Shah, Rajiv Ratn
author_facet Tank, Chayan
Mehta, Shaina
Pol, Sarthak
Katoch, Vinayak
Anand, Avinash
Jaiswal, Raj
Shah, Rajiv Ratn
contents In recent times, more and more people are posting about their mental states across various social media platforms. Leveraging this data, AI-based systems can be developed that help in assessing the mental health of individuals, such as suicide risk. This paper is a study done on suicidal risk assessments using Reddit data leveraging Base language models to identify patterns from social media posts. We have demonstrated that using smaller language models, i.e., less than 500M parameters, can also be effective in contrast to LLMs with greater than 500M parameters. We propose Su-RoBERTa, a fine-tuned RoBERTa on suicide risk prediction task that utilized both the labeled and unlabeled Reddit data and tackled class imbalance by data augmentation using GPT-2 model. Our Su-RoBERTa model attained a 69.84% weighted F1 score during the Final evaluation. This paper demonstrates the effectiveness of Base language models for the analysis of the risk factors related to mental health with an efficient computation pipeline
format Preprint
id arxiv_https___arxiv_org_abs_2412_01353
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Su-RoBERTa: A Semi-supervised Approach to Predicting Suicide Risk through Social Media using Base Language Models
Tank, Chayan
Mehta, Shaina
Pol, Sarthak
Katoch, Vinayak
Anand, Avinash
Jaiswal, Raj
Shah, Rajiv Ratn
Human-Computer Interaction
Artificial Intelligence
Social and Information Networks
In recent times, more and more people are posting about their mental states across various social media platforms. Leveraging this data, AI-based systems can be developed that help in assessing the mental health of individuals, such as suicide risk. This paper is a study done on suicidal risk assessments using Reddit data leveraging Base language models to identify patterns from social media posts. We have demonstrated that using smaller language models, i.e., less than 500M parameters, can also be effective in contrast to LLMs with greater than 500M parameters. We propose Su-RoBERTa, a fine-tuned RoBERTa on suicide risk prediction task that utilized both the labeled and unlabeled Reddit data and tackled class imbalance by data augmentation using GPT-2 model. Our Su-RoBERTa model attained a 69.84% weighted F1 score during the Final evaluation. This paper demonstrates the effectiveness of Base language models for the analysis of the risk factors related to mental health with an efficient computation pipeline
title Su-RoBERTa: A Semi-supervised Approach to Predicting Suicide Risk through Social Media using Base Language Models
topic Human-Computer Interaction
Artificial Intelligence
Social and Information Networks
url https://arxiv.org/abs/2412.01353