GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zaratiana, Urchade, Pasternak, Gil, Boyd, Oliver, Hurn-Maloney, George, Lewis, Ash
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915408374661120
author Zaratiana, Urchade
Pasternak, Gil
Boyd, Oliver
Hurn-Maloney, George
Lewis, Ash
author_facet Zaratiana, Urchade
Pasternak, Gil
Boyd, Oliver
Hurn-Maloney, George
Lewis, Ash
contents Information extraction (IE) is fundamental to numerous NLP applications, yet existing solutions often require specialized models for different tasks or rely on computationally expensive large language models. We present GLiNER2, a unified framework that enhances the original GLiNER architecture to support named entity recognition, text classification, and hierarchical structured data extraction within a single efficient model. Built pretrained transformer encoder architecture, GLiNER2 maintains CPU efficiency and compact size while introducing multi-task composition through an intuitive schema-based interface. Our experiments demonstrate competitive performance across extraction and classification tasks with substantial improvements in deployment accessibility compared to LLM-based alternatives. We release GLiNER2 as an open-source pip-installable library with pre-trained models and documentation at https://github.com/fastino-ai/GLiNER2.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18546
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface
Zaratiana, Urchade
Pasternak, Gil
Boyd, Oliver
Hurn-Maloney, George
Lewis, Ash
Computation and Language
Artificial Intelligence
Information extraction (IE) is fundamental to numerous NLP applications, yet existing solutions often require specialized models for different tasks or rely on computationally expensive large language models. We present GLiNER2, a unified framework that enhances the original GLiNER architecture to support named entity recognition, text classification, and hierarchical structured data extraction within a single efficient model. Built pretrained transformer encoder architecture, GLiNER2 maintains CPU efficiency and compact size while introducing multi-task composition through an intuitive schema-based interface. Our experiments demonstrate competitive performance across extraction and classification tasks with substantial improvements in deployment accessibility compared to LLM-based alternatives. We release GLiNER2 as an open-source pip-installable library with pre-trained models and documentation at https://github.com/fastino-ai/GLiNER2.
title GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2507.18546