Clustering Document Parts: Detecting and Characterizing Influence Campaigns from Documents

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Zhengxiang, Rambow, Owen
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914773289926656
author Wang, Zhengxiang
Rambow, Owen
author_facet Wang, Zhengxiang
Rambow, Owen
contents We propose a novel clustering pipeline to detect and characterize influence campaigns from documents. This approach clusters parts of document, detects clusters that likely reflect an influence campaign, and then identifies documents linked to an influence campaign via their association with the high-influence clusters. Our approach outperforms both the direct document-level classification and the direct document-level clustering approach in predicting if a document is part of an influence campaign. We propose various novel techniques to enhance our pipeline, including using an existing event factuality prediction system to obtain document parts, and aggregating multiple clustering experiments to improve the performance of both cluster and document classification. Classifying documents after clustering not only accurately extracts the parts of the documents that are relevant to influence campaigns, but also captures influence campaigns as a coordinated and holistic phenomenon. Our approach makes possible more fine-grained and interpretable characterizations of influence campaigns from documents.
format Preprint
id arxiv_https___arxiv_org_abs_2402_17151
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Clustering Document Parts: Detecting and Characterizing Influence Campaigns from Documents
Wang, Zhengxiang
Rambow, Owen
Computation and Language
We propose a novel clustering pipeline to detect and characterize influence campaigns from documents. This approach clusters parts of document, detects clusters that likely reflect an influence campaign, and then identifies documents linked to an influence campaign via their association with the high-influence clusters. Our approach outperforms both the direct document-level classification and the direct document-level clustering approach in predicting if a document is part of an influence campaign. We propose various novel techniques to enhance our pipeline, including using an existing event factuality prediction system to obtain document parts, and aggregating multiple clustering experiments to improve the performance of both cluster and document classification. Classifying documents after clustering not only accurately extracts the parts of the documents that are relevant to influence campaigns, but also captures influence campaigns as a coordinated and holistic phenomenon. Our approach makes possible more fine-grained and interpretable characterizations of influence campaigns from documents.
title Clustering Document Parts: Detecting and Characterizing Influence Campaigns from Documents
topic Computation and Language
url https://arxiv.org/abs/2402.17151