SciDef: Automating Definition Extraction from Academic Literature with Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kučera, Filip, Mandl, Christoph, Echizen, Isao, Timofte, Radu, Spinde, Timo
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915777642233856
author Kučera, Filip
Mandl, Christoph
Echizen, Isao
Timofte, Radu
Spinde, Timo
author_facet Kučera, Filip
Mandl, Christoph
Echizen, Isao
Timofte, Radu
Spinde, Timo
contents Definitions are the foundation for any scientific work, but with a significant increase in publication numbers, gathering definitions relevant to any keyword has become challenging. We therefore introduce SciDef, an LLM-based pipeline for automated definition extraction. We test SciDef on DefExtra & DefSim, novel datasets of human-extracted definitions and definition-pairs' similarity, respectively. Evaluating 16 language models across prompting strategies, we demonstrate that multi-step and DSPy-optimized prompting improve extraction performance. To evaluate extraction, we test various metrics and show that an NLI-based method yields the most reliable results. We show that LLMs are largely able to extract definitions from scientific literature (86.4% of definitions from our test-set); yet future work should focus not just on finding definitions, but on identifying relevant ones, as models tend to over-generate them. Code & datasets are available at https://github.com/Media-Bias-Group/SciDef.
format Preprint
id arxiv_https___arxiv_org_abs_2602_05413
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SciDef: Automating Definition Extraction from Academic Literature with Large Language Models
Kučera, Filip
Mandl, Christoph
Echizen, Isao
Timofte, Radu
Spinde, Timo
Information Retrieval
Computation and Language
Definitions are the foundation for any scientific work, but with a significant increase in publication numbers, gathering definitions relevant to any keyword has become challenging. We therefore introduce SciDef, an LLM-based pipeline for automated definition extraction. We test SciDef on DefExtra & DefSim, novel datasets of human-extracted definitions and definition-pairs' similarity, respectively. Evaluating 16 language models across prompting strategies, we demonstrate that multi-step and DSPy-optimized prompting improve extraction performance. To evaluate extraction, we test various metrics and show that an NLI-based method yields the most reliable results. We show that LLMs are largely able to extract definitions from scientific literature (86.4% of definitions from our test-set); yet future work should focus not just on finding definitions, but on identifying relevant ones, as models tend to over-generate them. Code & datasets are available at https://github.com/Media-Bias-Group/SciDef.
title SciDef: Automating Definition Extraction from Academic Literature with Large Language Models
topic Information Retrieval
Computation and Language
url https://arxiv.org/abs/2602.05413