Predicting the Geolocation of Tweets Using transformer models on Customized Data

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lutsai, Kateryna, Lampert, Christoph H.
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916559401779200
author Lutsai, Kateryna
Lampert, Christoph H.
author_facet Lutsai, Kateryna
Lampert, Christoph H.
contents This research is aimed to solve the tweet/user geolocation prediction task and provide a flexible methodology for the geotagging of textual big data. The suggested approach implements neural networks for natural language processing (NLP) to estimate the location as coordinate pairs (longitude, latitude) and two-dimensional Gaussian Mixture Models (GMMs). The scope of proposed models has been finetuned on a Twitter dataset using pretrained Bidirectional Encoder Representations from Transformers (BERT) as base models. Performance metrics show a median error of fewer than 30 km on a worldwide-level, and fewer than 15 km on the US-level datasets for the models trained and evaluated on text features of tweets' content and metadata context. Our source code and data are available at https://github.com/K4TEL/geo-twitter.git
format Preprint
id arxiv_https___arxiv_org_abs_2303_07865
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Predicting the Geolocation of Tweets Using transformer models on Customized Data
Lutsai, Kateryna
Lampert, Christoph H.
Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
68T50
I.2.7
This research is aimed to solve the tweet/user geolocation prediction task and provide a flexible methodology for the geotagging of textual big data. The suggested approach implements neural networks for natural language processing (NLP) to estimate the location as coordinate pairs (longitude, latitude) and two-dimensional Gaussian Mixture Models (GMMs). The scope of proposed models has been finetuned on a Twitter dataset using pretrained Bidirectional Encoder Representations from Transformers (BERT) as base models. Performance metrics show a median error of fewer than 30 km on a worldwide-level, and fewer than 15 km on the US-level datasets for the models trained and evaluated on text features of tweets' content and metadata context. Our source code and data are available at https://github.com/K4TEL/geo-twitter.git
title Predicting the Geolocation of Tweets Using transformer models on Customized Data
topic Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
68T50
I.2.7
url https://arxiv.org/abs/2303.07865