Stress Detection on Code-Mixed Texts in Dravidian Languages using Machine Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ramos, L., Shahiki-Tash, M., Ahani, Z., Eponon, A., Kolesnikova, O., Calvo, H.
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916428730335232
author Ramos, L.
Shahiki-Tash, M.
Ahani, Z.
Eponon, A.
Kolesnikova, O.
Calvo, H.
author_facet Ramos, L.
Shahiki-Tash, M.
Ahani, Z.
Eponon, A.
Kolesnikova, O.
Calvo, H.
contents Stress is a common feeling in daily life, but it can affect mental well-being in some situations, the development of robust detection models is imperative. This study introduces a methodical approach to the stress identification in code-mixed texts for Dravidian languages. The challenge encompassed two datasets, targeting Tamil and Telugu languages respectively. This proposal underscores the importance of using uncleaned text as a benchmark to refine future classification methodologies, incorporating diverse preprocessing techniques. Random Forest algorithm was used, featuring three textual representations: TF-IDF, Uni-grams of words, and a composite of (1+2+3)-Grams of characters. The approach achieved a good performance for both linguistic categories, achieving a Macro F1-score of 0.734 in Tamil and 0.727 in Telugu, overpassing results achieved with different complex techniques such as FastText and Transformer models. The results underscore the value of uncleaned data for mental state detection and the challenges classifying code-mixed texts for stress, indicating the potential for improved performance through cleaning data, other preprocessing techniques, or more complex models.
format Preprint
id arxiv_https___arxiv_org_abs_2410_06428
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Stress Detection on Code-Mixed Texts in Dravidian Languages using Machine Learning
Ramos, L.
Shahiki-Tash, M.
Ahani, Z.
Eponon, A.
Kolesnikova, O.
Calvo, H.
Computation and Language
Artificial Intelligence
Human-Computer Interaction
Machine Learning
Stress is a common feeling in daily life, but it can affect mental well-being in some situations, the development of robust detection models is imperative. This study introduces a methodical approach to the stress identification in code-mixed texts for Dravidian languages. The challenge encompassed two datasets, targeting Tamil and Telugu languages respectively. This proposal underscores the importance of using uncleaned text as a benchmark to refine future classification methodologies, incorporating diverse preprocessing techniques. Random Forest algorithm was used, featuring three textual representations: TF-IDF, Uni-grams of words, and a composite of (1+2+3)-Grams of characters. The approach achieved a good performance for both linguistic categories, achieving a Macro F1-score of 0.734 in Tamil and 0.727 in Telugu, overpassing results achieved with different complex techniques such as FastText and Transformer models. The results underscore the value of uncleaned data for mental state detection and the challenges classifying code-mixed texts for stress, indicating the potential for improved performance through cleaning data, other preprocessing techniques, or more complex models.
title Stress Detection on Code-Mixed Texts in Dravidian Languages using Machine Learning
topic Computation and Language
Artificial Intelligence
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2410.06428