GLiClass: Generalist Lightweight Model for Sequence Classification Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Stepanov, Ihor, Shtopko, Mykhailo, Vodianytskyi, Dmytro, Lukashov, Oleksandr, Yavorskyi, Alexander, Yaroshenko, Mykyta
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916891019182080
author Stepanov, Ihor
Shtopko, Mykhailo
Vodianytskyi, Dmytro
Lukashov, Oleksandr
Yavorskyi, Alexander
Yaroshenko, Mykyta
author_facet Stepanov, Ihor
Shtopko, Mykhailo
Vodianytskyi, Dmytro
Lukashov, Oleksandr
Yavorskyi, Alexander
Yaroshenko, Mykyta
contents Classification is one of the most widespread tasks in AI applications, serving often as the first step in filtering, sorting, and categorizing data. Since modern AI systems must handle large volumes of input data and early pipeline stages can propagate errors downstream, achieving high efficiency and accuracy is critical. Moreover, classification requirements can change dynamically based on user needs, necessitating models with strong zero-shot capabilities. While generative LLMs have become mainstream for zero-shot classification due to their versatility, they suffer from inconsistent instruction following and computational inefficiency. Cross-encoders, commonly used as rerankers in RAG pipelines, face a different bottleneck: they must process text-label pairs sequentially, significantly reducing efficiency with large label sets. Embedding-based approaches offer good efficiency but struggle with complex scenarios involving logical and semantic constraints. We propose GLiClass, a novel method that adapts the GLiNER architecture for sequence classification tasks. Our approach achieves strong accuracy and efficiency comparable to embedding-based methods, while maintaining the flexibility needed for zero-shot and few-shot learning scenarios. Additionally, we adapted proximal policy optimization (PPO) for multi-label text classification, enabling training classifiers in data-sparse conditions or from human feedback.
format Preprint
id arxiv_https___arxiv_org_abs_2508_07662
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GLiClass: Generalist Lightweight Model for Sequence Classification Tasks
Stepanov, Ihor
Shtopko, Mykhailo
Vodianytskyi, Dmytro
Lukashov, Oleksandr
Yavorskyi, Alexander
Yaroshenko, Mykyta
Machine Learning
Artificial Intelligence
Computation and Language
Classification is one of the most widespread tasks in AI applications, serving often as the first step in filtering, sorting, and categorizing data. Since modern AI systems must handle large volumes of input data and early pipeline stages can propagate errors downstream, achieving high efficiency and accuracy is critical. Moreover, classification requirements can change dynamically based on user needs, necessitating models with strong zero-shot capabilities. While generative LLMs have become mainstream for zero-shot classification due to their versatility, they suffer from inconsistent instruction following and computational inefficiency. Cross-encoders, commonly used as rerankers in RAG pipelines, face a different bottleneck: they must process text-label pairs sequentially, significantly reducing efficiency with large label sets. Embedding-based approaches offer good efficiency but struggle with complex scenarios involving logical and semantic constraints. We propose GLiClass, a novel method that adapts the GLiNER architecture for sequence classification tasks. Our approach achieves strong accuracy and efficiency comparable to embedding-based methods, while maintaining the flexibility needed for zero-shot and few-shot learning scenarios. Additionally, we adapted proximal policy optimization (PPO) for multi-label text classification, enabling training classifiers in data-sparse conditions or from human feedback.
title GLiClass: Generalist Lightweight Model for Sequence Classification Tasks
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.07662