Streamlining Social Media Information Retrieval for COVID-19 Research with Deep Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Hua, Yining, Wu, Jiageng, Lin, Shixu, Li, Minghui, Zhang, Yujie, Foer, Dinah, Wang, Siwen, Zhou, Peilin, Yang, Jie, Zhou, Li
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914716355395584
author Hua, Yining
Wu, Jiageng
Lin, Shixu
Li, Minghui
Zhang, Yujie
Foer, Dinah
Wang, Siwen
Zhou, Peilin
Yang, Jie
Zhou, Li
author_facet Hua, Yining
Wu, Jiageng
Lin, Shixu
Li, Minghui
Zhang, Yujie
Foer, Dinah
Wang, Siwen
Zhou, Peilin
Yang, Jie
Zhou, Li
contents Objective: Social media-based public health research is crucial for epidemic surveillance, but most studies identify relevant corpora with keyword-matching. This study develops a system to streamline the process of curating colloquial medical dictionaries. We demonstrate the pipeline by curating a UMLS-colloquial symptom dictionary from COVID-19-related tweets as proof of concept. Methods: COVID-19-related tweets from February 1, 2020, to April 30, 2022 were used. The pipeline includes three modules: a named entity recognition module to detect symptoms in tweets; an entity normalization module to aggregate detected entities; and a mapping module that iteratively maps entities to Unified Medical Language System concepts. A random 500 entity sample were drawn from the final dictionary for accuracy validation. Additionally, we conducted a symptom frequency distribution analysis to compare our dictionary to a pre-defined lexicon from previous research. Results: We identified 498,480 unique symptom entity expressions from the tweets. Pre-processing reduces the number to 18,226. The final dictionary contains 38,175 unique expressions of symptoms that can be mapped to 966 UMLS concepts (accuracy = 95%). Symptom distribution analysis found that our dictionary detects more symptoms and is effective at identifying psychiatric disorders like anxiety and depression, often missed by pre-defined lexicons. Conclusions: This study advances public health research by implementing a novel, systematic pipeline for curating symptom lexicons from social media data. The final lexicon's high accuracy, validated by medical professionals, underscores the potential of this methodology to reliably interpret and categorize vast amounts of unstructured social media data into actionable medical insights across diverse linguistic and regional landscapes.
format Preprint
id arxiv_https___arxiv_org_abs_2306_16001
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Streamlining Social Media Information Retrieval for COVID-19 Research with Deep Learning
Hua, Yining
Wu, Jiageng
Lin, Shixu
Li, Minghui
Zhang, Yujie
Foer, Dinah
Wang, Siwen
Zhou, Peilin
Yang, Jie
Zhou, Li
Computation and Language
Artificial Intelligence
Information Retrieval
Objective: Social media-based public health research is crucial for epidemic surveillance, but most studies identify relevant corpora with keyword-matching. This study develops a system to streamline the process of curating colloquial medical dictionaries. We demonstrate the pipeline by curating a UMLS-colloquial symptom dictionary from COVID-19-related tweets as proof of concept. Methods: COVID-19-related tweets from February 1, 2020, to April 30, 2022 were used. The pipeline includes three modules: a named entity recognition module to detect symptoms in tweets; an entity normalization module to aggregate detected entities; and a mapping module that iteratively maps entities to Unified Medical Language System concepts. A random 500 entity sample were drawn from the final dictionary for accuracy validation. Additionally, we conducted a symptom frequency distribution analysis to compare our dictionary to a pre-defined lexicon from previous research. Results: We identified 498,480 unique symptom entity expressions from the tweets. Pre-processing reduces the number to 18,226. The final dictionary contains 38,175 unique expressions of symptoms that can be mapped to 966 UMLS concepts (accuracy = 95%). Symptom distribution analysis found that our dictionary detects more symptoms and is effective at identifying psychiatric disorders like anxiety and depression, often missed by pre-defined lexicons. Conclusions: This study advances public health research by implementing a novel, systematic pipeline for curating symptom lexicons from social media data. The final lexicon's high accuracy, validated by medical professionals, underscores the potential of this methodology to reliably interpret and categorize vast amounts of unstructured social media data into actionable medical insights across diverse linguistic and regional landscapes.
title Streamlining Social Media Information Retrieval for COVID-19 Research with Deep Learning
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2306.16001