ThangDLU at #SMM4H 2024: Encoder-decoder models for classifying text data on social disorders in children and adolescents

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ta, Hoang-Thang, Rahman, Abu Bakar Siddiqur, Najjar, Lotfollah, Gelbukh, Alexander
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916230025183232
author Ta, Hoang-Thang
Rahman, Abu Bakar Siddiqur
Najjar, Lotfollah
Gelbukh, Alexander
author_facet Ta, Hoang-Thang
Rahman, Abu Bakar Siddiqur
Najjar, Lotfollah
Gelbukh, Alexander
contents This paper describes our participation in Task 3 and Task 5 of the #SMM4H (Social Media Mining for Health) 2024 Workshop, explicitly targeting the classification challenges within tweet data. Task 3 is a multi-class classification task centered on tweets discussing the impact of outdoor environments on symptoms of social anxiety. Task 5 involves a binary classification task focusing on tweets reporting medical disorders in children. We applied transfer learning from pre-trained encoder-decoder models such as BART-base and T5-small to identify the labels of a set of given tweets. We also presented some data augmentation methods to see their impact on the model performance. Finally, the systems obtained the best F1 score of 0.627 in Task 3 and the best F1 score of 0.841 in Task 5.
format Preprint
id arxiv_https___arxiv_org_abs_2404_19714
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ThangDLU at #SMM4H 2024: Encoder-decoder models for classifying text data on social disorders in children and adolescents
Ta, Hoang-Thang
Rahman, Abu Bakar Siddiqur
Najjar, Lotfollah
Gelbukh, Alexander
Computation and Language
This paper describes our participation in Task 3 and Task 5 of the #SMM4H (Social Media Mining for Health) 2024 Workshop, explicitly targeting the classification challenges within tweet data. Task 3 is a multi-class classification task centered on tweets discussing the impact of outdoor environments on symptoms of social anxiety. Task 5 involves a binary classification task focusing on tweets reporting medical disorders in children. We applied transfer learning from pre-trained encoder-decoder models such as BART-base and T5-small to identify the labels of a set of given tweets. We also presented some data augmentation methods to see their impact on the model performance. Finally, the systems obtained the best F1 score of 0.627 in Task 3 and the best F1 score of 0.841 in Task 5.
title ThangDLU at #SMM4H 2024: Encoder-decoder models for classifying text data on social disorders in children and adolescents
topic Computation and Language
url https://arxiv.org/abs/2404.19714