Large Language Models for Detection of Life-Threatening Texts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Thanh Thi, Wilson, Campbell, Dalins, Janis
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909647025209344
author Nguyen, Thanh Thi
Wilson, Campbell
Dalins, Janis
author_facet Nguyen, Thanh Thi
Wilson, Campbell
Dalins, Janis
contents Detecting life-threatening language is essential for safeguarding individuals in distress, promoting mental health and well-being, and preventing potential harm and loss of life. This paper presents an effective approach to identifying life-threatening texts using large language models (LLMs) and compares them with traditional methods such as bag of words, word embedding, topic modeling, and Bidirectional Encoder Representations from Transformers. We fine-tune three open-source LLMs including Gemma, Mistral, and Llama-2 using their 7B parameter variants on different datasets, which are constructed with class balance, imbalance, and extreme imbalance scenarios. Experimental results demonstrate a strong performance of LLMs against traditional methods. More specifically, Mistral and Llama-2 models are top performers in both balanced and imbalanced data scenarios while Gemma is slightly behind. We employ the upsampling technique to deal with the imbalanced data scenarios and demonstrate that while this method benefits traditional approaches, it does not have as much impact on LLMs. This study demonstrates a great potential of LLMs for real-world life-threatening language detection problems.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10687
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Large Language Models for Detection of Life-Threatening Texts
Nguyen, Thanh Thi
Wilson, Campbell
Dalins, Janis
Computation and Language
Artificial Intelligence
Machine Learning
Detecting life-threatening language is essential for safeguarding individuals in distress, promoting mental health and well-being, and preventing potential harm and loss of life. This paper presents an effective approach to identifying life-threatening texts using large language models (LLMs) and compares them with traditional methods such as bag of words, word embedding, topic modeling, and Bidirectional Encoder Representations from Transformers. We fine-tune three open-source LLMs including Gemma, Mistral, and Llama-2 using their 7B parameter variants on different datasets, which are constructed with class balance, imbalance, and extreme imbalance scenarios. Experimental results demonstrate a strong performance of LLMs against traditional methods. More specifically, Mistral and Llama-2 models are top performers in both balanced and imbalanced data scenarios while Gemma is slightly behind. We employ the upsampling technique to deal with the imbalanced data scenarios and demonstrate that while this method benefits traditional approaches, it does not have as much impact on LLMs. This study demonstrates a great potential of LLMs for real-world life-threatening language detection problems.
title Large Language Models for Detection of Life-Threatening Texts
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.10687