Leveraging Large Language Models for Generating Research Topic Ontologies: A Multi-Disciplinary Study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Aggarwal, Tanay, Salatino, Angelo, Osborne, Francesco, Motta, Enrico
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909757859692544
author Aggarwal, Tanay
Salatino, Angelo
Osborne, Francesco
Motta, Enrico
author_facet Aggarwal, Tanay
Salatino, Angelo
Osborne, Francesco
Motta, Enrico
contents Ontologies and taxonomies of research fields are critical for managing and organising scientific knowledge, as they facilitate efficient classification, dissemination and retrieval of information. However, the creation and maintenance of such ontologies are expensive and time-consuming tasks, usually requiring the coordinated effort of multiple domain experts. Consequently, ontologies in this space often exhibit uneven coverage across different disciplines, limited inter-domain connectivity, and infrequent updating cycles. In this study, we investigate the capability of several large language models to identify semantic relationships among research topics within three academic domains: biomedicine, physics, and engineering. The models were evaluated under three distinct conditions: zero-shot prompting, chain-of-thought prompting, and fine-tuning on existing ontologies. Additionally, we assessed the cross-domain transferability of fine-tuned models by measuring their performance when trained in one domain and subsequently applied to a different one. To support this analysis, we introduce PEM-Rel-8K, a novel dataset consisting of over 8,000 relationships extracted from the most widely adopted taxonomies in the three disciplines considered in this study: MeSH, PhySH, and IEEE. Our experiments demonstrate that fine-tuning LLMs on PEM-Rel-8K yields excellent performance across all disciplines.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20693
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging Large Language Models for Generating Research Topic Ontologies: A Multi-Disciplinary Study
Aggarwal, Tanay
Salatino, Angelo
Osborne, Francesco
Motta, Enrico
Digital Libraries
Computation and Language
Ontologies and taxonomies of research fields are critical for managing and organising scientific knowledge, as they facilitate efficient classification, dissemination and retrieval of information. However, the creation and maintenance of such ontologies are expensive and time-consuming tasks, usually requiring the coordinated effort of multiple domain experts. Consequently, ontologies in this space often exhibit uneven coverage across different disciplines, limited inter-domain connectivity, and infrequent updating cycles. In this study, we investigate the capability of several large language models to identify semantic relationships among research topics within three academic domains: biomedicine, physics, and engineering. The models were evaluated under three distinct conditions: zero-shot prompting, chain-of-thought prompting, and fine-tuning on existing ontologies. Additionally, we assessed the cross-domain transferability of fine-tuned models by measuring their performance when trained in one domain and subsequently applied to a different one. To support this analysis, we introduce PEM-Rel-8K, a novel dataset consisting of over 8,000 relationships extracted from the most widely adopted taxonomies in the three disciplines considered in this study: MeSH, PhySH, and IEEE. Our experiments demonstrate that fine-tuning LLMs on PEM-Rel-8K yields excellent performance across all disciplines.
title Leveraging Large Language Models for Generating Research Topic Ontologies: A Multi-Disciplinary Study
topic Digital Libraries
Computation and Language
url https://arxiv.org/abs/2508.20693