Emoji Retrieval from Gibberish or Garbled Social Media Text: A Novel Methodology and A Case Study

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cui, Shuqi, Thakur, Nirmalya, Poon, Audrey
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909440009043968
author Cui, Shuqi
Thakur, Nirmalya
Poon, Audrey
author_facet Cui, Shuqi
Thakur, Nirmalya
Poon, Audrey
contents Emojis are widely used across social media platforms but are often lost in noisy or garbled text, posing challenges for data analysis and machine learning. Conventional preprocessing approaches recommend removing such text, risking the loss of emojis and their contextual meaning. This paper proposes a three-step reverse-engineering methodology to retrieve emojis from garbled text in social media posts. The methodology also identifies reasons for the generation of such text during social media data mining. To evaluate its effectiveness, the approach was applied to 509,248 Tweets about the Mpox outbreak, a dataset referenced in about 30 prior works that failed to retrieve emojis from garbled text. Our method retrieved 157,748 emojis from 76,914 Tweets. Improvements in text readability and coherence were demonstrated through metrics such as Flesch Reading Ease, Flesch-Kincaid Grade Level, Coleman-Liau Index, Automated Readability Index, Dale-Chall Readability Score, Text Standard, and Reading Time. Additionally, the frequency of individual emojis and their patterns of usage in these Tweets were analyzed, and the results are presented.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18046
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Emoji Retrieval from Gibberish or Garbled Social Media Text: A Novel Methodology and A Case Study
Cui, Shuqi
Thakur, Nirmalya
Poon, Audrey
Social and Information Networks
Artificial Intelligence
Computation and Language
Computers and Society
Machine Learning
I.2.7; I.2.8; I.5.4; K.4.2; H.2.8; I.2.6
Emojis are widely used across social media platforms but are often lost in noisy or garbled text, posing challenges for data analysis and machine learning. Conventional preprocessing approaches recommend removing such text, risking the loss of emojis and their contextual meaning. This paper proposes a three-step reverse-engineering methodology to retrieve emojis from garbled text in social media posts. The methodology also identifies reasons for the generation of such text during social media data mining. To evaluate its effectiveness, the approach was applied to 509,248 Tweets about the Mpox outbreak, a dataset referenced in about 30 prior works that failed to retrieve emojis from garbled text. Our method retrieved 157,748 emojis from 76,914 Tweets. Improvements in text readability and coherence were demonstrated through metrics such as Flesch Reading Ease, Flesch-Kincaid Grade Level, Coleman-Liau Index, Automated Readability Index, Dale-Chall Readability Score, Text Standard, and Reading Time. Additionally, the frequency of individual emojis and their patterns of usage in these Tweets were analyzed, and the results are presented.
title Emoji Retrieval from Gibberish or Garbled Social Media Text: A Novel Methodology and A Case Study
topic Social and Information Networks
Artificial Intelligence
Computation and Language
Computers and Society
Machine Learning
I.2.7; I.2.8; I.5.4; K.4.2; H.2.8; I.2.6
url https://arxiv.org/abs/2412.18046