MeXtract: Light-Weight Metadata Extraction from Scientific Papers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alyafeai, Zaid, Al-Shaibani, Maged S., Ghanem, Bernard
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914080596426752
author Alyafeai, Zaid
Al-Shaibani, Maged S.
Ghanem, Bernard
author_facet Alyafeai, Zaid
Al-Shaibani, Maged S.
Ghanem, Bernard
contents Metadata plays a critical role in indexing, documenting, and analyzing scientific literature, yet extracting it accurately and efficiently remains a challenging task. Traditional approaches often rely on rule-based or task-specific models, which struggle to generalize across domains and schema variations. In this paper, we present MeXtract, a family of lightweight language models designed for metadata extraction from scientific papers. The models, ranging from 0.5B to 3B parameters, are built by fine-tuning Qwen 2.5 counterparts. In their size family, MeXtract achieves state-of-the-art performance on metadata extraction on the MOLE benchmark. To further support evaluation, we extend the MOLE benchmark to incorporate model-specific metadata, providing an out-of-domain challenging subset. Our experiments show that fine-tuning on a given schema not only yields high accuracy but also transfers effectively to unseen schemas, demonstrating the robustness and adaptability of our approach. We release all the code, datasets, and models openly for the research community.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06889
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MeXtract: Light-Weight Metadata Extraction from Scientific Papers
Alyafeai, Zaid
Al-Shaibani, Maged S.
Ghanem, Bernard
Computation and Language
Metadata plays a critical role in indexing, documenting, and analyzing scientific literature, yet extracting it accurately and efficiently remains a challenging task. Traditional approaches often rely on rule-based or task-specific models, which struggle to generalize across domains and schema variations. In this paper, we present MeXtract, a family of lightweight language models designed for metadata extraction from scientific papers. The models, ranging from 0.5B to 3B parameters, are built by fine-tuning Qwen 2.5 counterparts. In their size family, MeXtract achieves state-of-the-art performance on metadata extraction on the MOLE benchmark. To further support evaluation, we extend the MOLE benchmark to incorporate model-specific metadata, providing an out-of-domain challenging subset. Our experiments show that fine-tuning on a given schema not only yields high accuracy but also transfers effectively to unseen schemas, demonstrating the robustness and adaptability of our approach. We release all the code, datasets, and models openly for the research community.
title MeXtract: Light-Weight Metadata Extraction from Scientific Papers
topic Computation and Language
url https://arxiv.org/abs/2510.06889