Ontology Population using LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Norouzi, Sanaz Saki, Barua, Adrita, Christou, Antrea, Gautam, Nikita, Eells, Andrew, Hitzler, Pascal, Shimizu, Cogan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917827732045824
author Norouzi, Sanaz Saki
Barua, Adrita
Christou, Antrea
Gautam, Nikita
Eells, Andrew
Hitzler, Pascal
Shimizu, Cogan
author_facet Norouzi, Sanaz Saki
Barua, Adrita
Christou, Antrea
Gautam, Nikita
Eells, Andrew
Hitzler, Pascal
Shimizu, Cogan
contents Knowledge graphs (KGs) are increasingly utilized for data integration, representation, and visualization. While KG population is critical, it is often costly, especially when data must be extracted from unstructured text in natural language, which presents challenges, such as ambiguity and complex interpretations. Large Language Models (LLMs) offer promising capabilities for such tasks, excelling in natural language understanding and content generation. However, their tendency to ``hallucinate'' can produce inaccurate outputs. Despite these limitations, LLMs offer rapid and scalable processing of natural language data, and with prompt engineering and fine-tuning, they can approximate human-level performance in extracting and structuring data for KGs. This study investigates LLM effectiveness for the KG population, focusing on the Enslaved.org Hub Ontology. In this paper, we report that compared to the ground truth, LLM's can extract ~90% of triples, when provided a modular ontology as guidance in the prompts.
format Preprint
id arxiv_https___arxiv_org_abs_2411_01612
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Ontology Population using LLMs
Norouzi, Sanaz Saki
Barua, Adrita
Christou, Antrea
Gautam, Nikita
Eells, Andrew
Hitzler, Pascal
Shimizu, Cogan
Artificial Intelligence
Computation and Language
Knowledge graphs (KGs) are increasingly utilized for data integration, representation, and visualization. While KG population is critical, it is often costly, especially when data must be extracted from unstructured text in natural language, which presents challenges, such as ambiguity and complex interpretations. Large Language Models (LLMs) offer promising capabilities for such tasks, excelling in natural language understanding and content generation. However, their tendency to ``hallucinate'' can produce inaccurate outputs. Despite these limitations, LLMs offer rapid and scalable processing of natural language data, and with prompt engineering and fine-tuning, they can approximate human-level performance in extracting and structuring data for KGs. This study investigates LLM effectiveness for the KG population, focusing on the Enslaved.org Hub Ontology. In this paper, we report that compared to the ground truth, LLM's can extract ~90% of triples, when provided a modular ontology as guidance in the prompts.
title Ontology Population using LLMs
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2411.01612