KnowledgeHub: An end-to-end Tool for Assisted Scientific Discovery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tanaka, Shinnosuke, Barry, James, Kuruvanthodi, Vishnudev, Moses, Movina, Giammona, Maxwell J., Herr, Nathan, Elkaref, Mohab, De Mel, Geeth
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913490051006464
author Tanaka, Shinnosuke
Barry, James
Kuruvanthodi, Vishnudev
Moses, Movina
Giammona, Maxwell J.
Herr, Nathan
Elkaref, Mohab
De Mel, Geeth
author_facet Tanaka, Shinnosuke
Barry, James
Kuruvanthodi, Vishnudev
Moses, Movina
Giammona, Maxwell J.
Herr, Nathan
Elkaref, Mohab
De Mel, Geeth
contents This paper describes the KnowledgeHub tool, a scientific literature Information Extraction (IE) and Question Answering (QA) pipeline. This is achieved by supporting the ingestion of PDF documents that are converted to text and structured representations. An ontology can then be constructed where a user defines the types of entities and relationships they want to capture. A browser-based annotation tool enables annotating the contents of the PDF documents according to the ontology. Named Entity Recognition (NER) and Relation Classification (RC) models can be trained on the resulting annotations and can be used to annotate the unannotated portion of the documents. A knowledge graph is constructed from these entity and relation triples which can be queried to obtain insights from the data. Furthermore, we integrate a suite of Large Language Models (LLMs) that can be used for QA and summarisation that is grounded in the included documents via a retrieval component. KnowledgeHub is a unique tool that supports annotation, IE and QA, which gives the user full insight into the knowledge discovery pipeline.
format Preprint
id arxiv_https___arxiv_org_abs_2406_00008
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle KnowledgeHub: An end-to-end Tool for Assisted Scientific Discovery
Tanaka, Shinnosuke
Barry, James
Kuruvanthodi, Vishnudev
Moses, Movina
Giammona, Maxwell J.
Herr, Nathan
Elkaref, Mohab
De Mel, Geeth
Information Retrieval
Artificial Intelligence
Computation and Language
Digital Libraries
This paper describes the KnowledgeHub tool, a scientific literature Information Extraction (IE) and Question Answering (QA) pipeline. This is achieved by supporting the ingestion of PDF documents that are converted to text and structured representations. An ontology can then be constructed where a user defines the types of entities and relationships they want to capture. A browser-based annotation tool enables annotating the contents of the PDF documents according to the ontology. Named Entity Recognition (NER) and Relation Classification (RC) models can be trained on the resulting annotations and can be used to annotate the unannotated portion of the documents. A knowledge graph is constructed from these entity and relation triples which can be queried to obtain insights from the data. Furthermore, we integrate a suite of Large Language Models (LLMs) that can be used for QA and summarisation that is grounded in the included documents via a retrieval component. KnowledgeHub is a unique tool that supports annotation, IE and QA, which gives the user full insight into the knowledge discovery pipeline.
title KnowledgeHub: An end-to-end Tool for Assisted Scientific Discovery
topic Information Retrieval
Artificial Intelligence
Computation and Language
Digital Libraries
url https://arxiv.org/abs/2406.00008