OAEI-LLM-T: A TBox Benchmark Dataset for Understanding Large Language Model Hallucinations in Ontology Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qiang, Zhangcheng, Taylor, Kerry, Wang, Weiqing, Jiang, Jing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909982426923008
author Qiang, Zhangcheng
Taylor, Kerry
Wang, Weiqing
Jiang, Jing
author_facet Qiang, Zhangcheng
Taylor, Kerry
Wang, Weiqing
Jiang, Jing
contents Hallucinations are often inevitable in downstream tasks using large language models (LLMs). To tackle the substantial challenge of addressing hallucinations for LLM-based ontology matching (OM) systems, we introduce a new benchmark dataset OAEI-LLM-T. The dataset evolves from seven TBox datasets in the Ontology Alignment Evaluation Initiative (OAEI), capturing hallucinations of ten different LLMs performing OM tasks. These OM-specific hallucinations are organised into two primary categories and six sub-categories. We showcase the usefulness of the dataset in constructing an LLM leaderboard for OM tasks and for fine-tuning LLMs used in OM tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21813
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OAEI-LLM-T: A TBox Benchmark Dataset for Understanding Large Language Model Hallucinations in Ontology Matching
Qiang, Zhangcheng
Taylor, Kerry
Wang, Weiqing
Jiang, Jing
Artificial Intelligence
Computation and Language
Information Retrieval
Machine Learning
Hallucinations are often inevitable in downstream tasks using large language models (LLMs). To tackle the substantial challenge of addressing hallucinations for LLM-based ontology matching (OM) systems, we introduce a new benchmark dataset OAEI-LLM-T. The dataset evolves from seven TBox datasets in the Ontology Alignment Evaluation Initiative (OAEI), capturing hallucinations of ten different LLMs performing OM tasks. These OM-specific hallucinations are organised into two primary categories and six sub-categories. We showcase the usefulness of the dataset in constructing an LLM leaderboard for OM tasks and for fine-tuning LLMs used in OM tasks.
title OAEI-LLM-T: A TBox Benchmark Dataset for Understanding Large Language Model Hallucinations in Ontology Matching
topic Artificial Intelligence
Computation and Language
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2503.21813