Modelling the Spread of New Information on X

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xu, Ziming, Zhou, Shi, Lampos, Vasileios, Cox, Ingemar J.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909992340160512
author Xu, Ziming
Zhou, Shi
Lampos, Vasileios
Cox, Ingemar J.
author_facet Xu, Ziming
Zhou, Shi
Lampos, Vasileios
Cox, Ingemar J.
contents There has been considerable interest in modelling the spread of information on X (formerly Twitter) using machine learning models. Here, we consider the problem of predicting the reposting of new information, i.e., when a user propagates information about a topic previously unseen by the user. In existing work, information and users are randomly assigned to a test or training set, ensuring that both sets are drawn from the same distribution. In the spread of new information, the problem becomes an out-of-distribution classification task. Our experimental results reveal that while existing algorithms, which predominantly use features derived from the content of posts, perform well when the training and test distributions are the same, they perform much worse when the test set is out-of-distribution, i.e., when the topic of the testing data is absent from the training data. We then show that if the post features are supplemented or replaced with features derived from user profiles and past behaviours, the out-of-distribution prediction is greatly improved, with the F1 score increasing from 0.117 to 0.705. Our experimental results suggest that a significant component of reposting behaviour for previously unseen topics can be predicted from user profiles and past behaviours, and is largely content-agnostic.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15370
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Modelling the Spread of New Information on X
Xu, Ziming
Zhou, Shi
Lampos, Vasileios
Cox, Ingemar J.
Social and Information Networks
There has been considerable interest in modelling the spread of information on X (formerly Twitter) using machine learning models. Here, we consider the problem of predicting the reposting of new information, i.e., when a user propagates information about a topic previously unseen by the user. In existing work, information and users are randomly assigned to a test or training set, ensuring that both sets are drawn from the same distribution. In the spread of new information, the problem becomes an out-of-distribution classification task. Our experimental results reveal that while existing algorithms, which predominantly use features derived from the content of posts, perform well when the training and test distributions are the same, they perform much worse when the test set is out-of-distribution, i.e., when the topic of the testing data is absent from the training data. We then show that if the post features are supplemented or replaced with features derived from user profiles and past behaviours, the out-of-distribution prediction is greatly improved, with the F1 score increasing from 0.117 to 0.705. Our experimental results suggest that a significant component of reposting behaviour for previously unseen topics can be predicted from user profiles and past behaviours, and is largely content-agnostic.
title Modelling the Spread of New Information on X
topic Social and Information Networks
url https://arxiv.org/abs/2505.15370