From low resource information extraction to identifying influential nodes in knowledge graphs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Cai, Erica, Simek, Olga, Miller, Benjamin A., Sullivan-Pao, Danielle, Young, Evan, Smith, Christopher L.
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910293118943232
author Cai, Erica
Simek, Olga
Miller, Benjamin A.
Sullivan-Pao, Danielle
Young, Evan
Smith, Christopher L.
author_facet Cai, Erica
Simek, Olga
Miller, Benjamin A.
Sullivan-Pao, Danielle
Young, Evan
Smith, Christopher L.
contents We propose a pipeline for identifying important entities from intelligence reports that constructs a knowledge graph, where nodes correspond to entities of fine-grained types (e.g. traffickers) extracted from the text and edges correspond to extracted relations between entities (e.g. cartel membership). The important entities in intelligence reports then map to central nodes in the knowledge graph. We introduce a novel method that extracts fine-grained entities in a few-shot setting (few labeled examples), given limited resources available to label the frequently changing entity types that intelligence analysts are interested in. It outperforms other state-of-the-art methods. Next, we identify challenges facing previous evaluations of zero-shot (no labeled examples) methods for extracting relations, affecting the step of populating edges. Finally, we explore the utility of the pipeline: given the goal of identifying important entities, we evaluate the impact of relation extraction errors on the identification of central nodes in several real and synthetic networks. The impact of these errors varies significantly by graph topology, suggesting that confidence in measurements based on automatically extracted relations should depend on observed network features.
format Preprint
id arxiv_https___arxiv_org_abs_2401_04915
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle From low resource information extraction to identifying influential nodes in knowledge graphs
Cai, Erica
Simek, Olga
Miller, Benjamin A.
Sullivan-Pao, Danielle
Young, Evan
Smith, Christopher L.
Social and Information Networks
We propose a pipeline for identifying important entities from intelligence reports that constructs a knowledge graph, where nodes correspond to entities of fine-grained types (e.g. traffickers) extracted from the text and edges correspond to extracted relations between entities (e.g. cartel membership). The important entities in intelligence reports then map to central nodes in the knowledge graph. We introduce a novel method that extracts fine-grained entities in a few-shot setting (few labeled examples), given limited resources available to label the frequently changing entity types that intelligence analysts are interested in. It outperforms other state-of-the-art methods. Next, we identify challenges facing previous evaluations of zero-shot (no labeled examples) methods for extracting relations, affecting the step of populating edges. Finally, we explore the utility of the pipeline: given the goal of identifying important entities, we evaluate the impact of relation extraction errors on the identification of central nodes in several real and synthetic networks. The impact of these errors varies significantly by graph topology, suggesting that confidence in measurements based on automatically extracted relations should depend on observed network features.
title From low resource information extraction to identifying influential nodes in knowledge graphs
topic Social and Information Networks
url https://arxiv.org/abs/2401.04915