SpannerLib: Embedding Declarative Information Extraction in an Imperative Workflow

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Light, Dean, Aiashy, Ahmad, Diab, Mahmoud, Nachmias, Daniel, Vansummeren, Stijn, Kimelfeld, Benny
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929484088737792
author Light, Dean
Aiashy, Ahmad
Diab, Mahmoud
Nachmias, Daniel
Vansummeren, Stijn
Kimelfeld, Benny
author_facet Light, Dean
Aiashy, Ahmad
Diab, Mahmoud
Nachmias, Daniel
Vansummeren, Stijn
Kimelfeld, Benny
contents Document spanners have been proposed as a formal framework for declarative Information Extraction (IE) from text, following IE products from the industry and academia. Over the past decade, the framework has been studied thoroughly in terms of expressive power, complexity, and the ability to naturally combine text analysis with relational querying. This demonstration presents SpannerLib a library for embedding document spanners in Python code. SpannerLib facilitates the development of IE programs by providing an implementation of Spannerlog (Datalog-based documentspanners) that interacts with the Python code in two directions: rules can be embedded inside Python, and they can invoke custom Python code (e.g., calls to ML-based NLP models) via user-defined functions. The demonstration scenarios showcase IE programs, with increasing levels of complexity, within Jupyter Notebook.
format Preprint
id arxiv_https___arxiv_org_abs_2409_01736
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SpannerLib: Embedding Declarative Information Extraction in an Imperative Workflow
Light, Dean
Aiashy, Ahmad
Diab, Mahmoud
Nachmias, Daniel
Vansummeren, Stijn
Kimelfeld, Benny
Databases
Information Retrieval
H.4
Document spanners have been proposed as a formal framework for declarative Information Extraction (IE) from text, following IE products from the industry and academia. Over the past decade, the framework has been studied thoroughly in terms of expressive power, complexity, and the ability to naturally combine text analysis with relational querying. This demonstration presents SpannerLib a library for embedding document spanners in Python code. SpannerLib facilitates the development of IE programs by providing an implementation of Spannerlog (Datalog-based documentspanners) that interacts with the Python code in two directions: rules can be embedded inside Python, and they can invoke custom Python code (e.g., calls to ML-based NLP models) via user-defined functions. The demonstration scenarios showcase IE programs, with increasing levels of complexity, within Jupyter Notebook.
title SpannerLib: Embedding Declarative Information Extraction in an Imperative Workflow
topic Databases
Information Retrieval
H.4
url https://arxiv.org/abs/2409.01736