ProUIE: A Macro-to-Micro Progressive Learning Method for LLM-based Universal Information Extraction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Wenda, Song, Zhigang, Nie, Shuai, Liu, Guangyao, Chen, Lisung, Yang, Binyu, Chen, Yaran, Zhou, Peng, Wang, Hongzhen, Liu, Yuchen, Hu, Wenyue, Xu, Jiaming, Shi, Runyu, Huang, Ying
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913025010696192
author Liu, Wenda
Song, Zhigang
Nie, Shuai
Liu, Guangyao
Chen, Lisung
Yang, Binyu
Chen, Yaran
Zhou, Peng
Wang, Hongzhen
Liu, Yuchen
Hu, Wenyue
Xu, Jiaming
Shi, Runyu
Huang, Ying
author_facet Liu, Wenda
Song, Zhigang
Nie, Shuai
Liu, Guangyao
Chen, Lisung
Yang, Binyu
Chen, Yaran
Zhou, Peng
Wang, Hongzhen
Liu, Yuchen
Hu, Wenyue
Xu, Jiaming
Shi, Runyu
Huang, Ying
contents LLM-based universal information extraction (UIE) methods often rely on additional information beyond the original training data, which increases training complexity yet often yields limited gains. To address this, we propose ProUIE, a Macro-to-Micro progressive learning approach that improves UIE without introducing any external information. ProUIE consists of three stages: (i) macro-level Complete Modeling (CM), which learns NER, RE, and EE along their intrinsic difficulty order on the full training data to build a unified extraction foundation, (ii) meso-level Streamlined Alignment (SA), which operates on sampled data with simplified target formats, streamlining and regularizing structured outputs to make them more concise and controllable, and (iii) micro-level Deep Exploration (DE), which applies GRPO with stepwise fine-grained rewards (SFR) over structural units to guide exploration and improve performance. Experiments on 36 public datasets show that ProUIE consistently improves unified extraction, outperforming strong instruction-tuned baselines on average for NER and RE while using a smaller backbone, and it further demonstrates clear gains in large-scale production-oriented information extraction.
format Preprint
id arxiv_https___arxiv_org_abs_2604_10633
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ProUIE: A Macro-to-Micro Progressive Learning Method for LLM-based Universal Information Extraction
Liu, Wenda
Song, Zhigang
Nie, Shuai
Liu, Guangyao
Chen, Lisung
Yang, Binyu
Chen, Yaran
Zhou, Peng
Wang, Hongzhen
Liu, Yuchen
Hu, Wenyue
Xu, Jiaming
Shi, Runyu
Huang, Ying
Computation and Language
LLM-based universal information extraction (UIE) methods often rely on additional information beyond the original training data, which increases training complexity yet often yields limited gains. To address this, we propose ProUIE, a Macro-to-Micro progressive learning approach that improves UIE without introducing any external information. ProUIE consists of three stages: (i) macro-level Complete Modeling (CM), which learns NER, RE, and EE along their intrinsic difficulty order on the full training data to build a unified extraction foundation, (ii) meso-level Streamlined Alignment (SA), which operates on sampled data with simplified target formats, streamlining and regularizing structured outputs to make them more concise and controllable, and (iii) micro-level Deep Exploration (DE), which applies GRPO with stepwise fine-grained rewards (SFR) over structural units to guide exploration and improve performance. Experiments on 36 public datasets show that ProUIE consistently improves unified extraction, outperforming strong instruction-tuned baselines on average for NER and RE while using a smaller backbone, and it further demonstrates clear gains in large-scale production-oriented information extraction.
title ProUIE: A Macro-to-Micro Progressive Learning Method for LLM-based Universal Information Extraction
topic Computation and Language
url https://arxiv.org/abs/2604.10633