SPEED++: A Multilingual Event Extraction Framework for Epidemic Prediction and Preparedness

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Parekh, Tanmay, Kwan, Jeffrey, Yu, Jiarui, Johri, Sparsh, Ahn, Hyosang, Muppalla, Sreya, Chang, Kai-Wei, Wang, Wei, Peng, Nanyun
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913562344030208
author Parekh, Tanmay
Kwan, Jeffrey
Yu, Jiarui
Johri, Sparsh
Ahn, Hyosang
Muppalla, Sreya
Chang, Kai-Wei
Wang, Wei
Peng, Nanyun
author_facet Parekh, Tanmay
Kwan, Jeffrey
Yu, Jiarui
Johri, Sparsh
Ahn, Hyosang
Muppalla, Sreya
Chang, Kai-Wei
Wang, Wei
Peng, Nanyun
contents Social media is often the first place where communities discuss the latest societal trends. Prior works have utilized this platform to extract epidemic-related information (e.g. infections, preventive measures) to provide early warnings for epidemic prediction. However, these works only focused on English posts, while epidemics can occur anywhere in the world, and early discussions are often in the local, non-English languages. In this work, we introduce the first multilingual Event Extraction (EE) framework SPEED++ for extracting epidemic event information for a wide range of diseases and languages. To this end, we extend a previous epidemic ontology with 20 argument roles; and curate our multilingual EE dataset SPEED++ comprising 5.1K tweets in four languages for four diseases. Annotating data in every language is infeasible; thus we develop zero-shot cross-lingual cross-disease models (i.e., training only on English COVID data) utilizing multilingual pre-training and show their efficacy in extracting epidemic-related events for 65 diverse languages across different diseases. Experiments demonstrate that our framework can provide epidemic warnings for COVID-19 in its earliest stages in Dec 2019 (3 weeks before global discussions) from Chinese Weibo posts without any training in Chinese. Furthermore, we exploit our framework's argument extraction capabilities to aggregate community epidemic discussions like symptoms and cure measures, aiding misinformation detection and public attention monitoring. Overall, we lay a strong foundation for multilingual epidemic preparedness.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18393
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SPEED++: A Multilingual Event Extraction Framework for Epidemic Prediction and Preparedness
Parekh, Tanmay
Kwan, Jeffrey
Yu, Jiarui
Johri, Sparsh
Ahn, Hyosang
Muppalla, Sreya
Chang, Kai-Wei
Wang, Wei
Peng, Nanyun
Computation and Language
Social and Information Networks
Social media is often the first place where communities discuss the latest societal trends. Prior works have utilized this platform to extract epidemic-related information (e.g. infections, preventive measures) to provide early warnings for epidemic prediction. However, these works only focused on English posts, while epidemics can occur anywhere in the world, and early discussions are often in the local, non-English languages. In this work, we introduce the first multilingual Event Extraction (EE) framework SPEED++ for extracting epidemic event information for a wide range of diseases and languages. To this end, we extend a previous epidemic ontology with 20 argument roles; and curate our multilingual EE dataset SPEED++ comprising 5.1K tweets in four languages for four diseases. Annotating data in every language is infeasible; thus we develop zero-shot cross-lingual cross-disease models (i.e., training only on English COVID data) utilizing multilingual pre-training and show their efficacy in extracting epidemic-related events for 65 diverse languages across different diseases. Experiments demonstrate that our framework can provide epidemic warnings for COVID-19 in its earliest stages in Dec 2019 (3 weeks before global discussions) from Chinese Weibo posts without any training in Chinese. Furthermore, we exploit our framework's argument extraction capabilities to aggregate community epidemic discussions like symptoms and cure measures, aiding misinformation detection and public attention monitoring. Overall, we lay a strong foundation for multilingual epidemic preparedness.
title SPEED++: A Multilingual Event Extraction Framework for Epidemic Prediction and Preparedness
topic Computation and Language
Social and Information Networks
url https://arxiv.org/abs/2410.18393