ReMatch: Retrieval Enhanced Schema Matching with LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sheetrit, Eitam, Brief, Menachem, Mishaeli, Moshik, Elisha, Oren
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914816190316544
author Sheetrit, Eitam
Brief, Menachem
Mishaeli, Moshik
Elisha, Oren
author_facet Sheetrit, Eitam
Brief, Menachem
Mishaeli, Moshik
Elisha, Oren
contents Schema matching is a crucial task in data integration, involving the alignment of a source schema with a target schema to establish correspondence between their elements. This task is challenging due to textual and semantic heterogeneity, as well as differences in schema sizes. Although machine-learning-based solutions have been explored in numerous studies, they often suffer from low accuracy, require manual mapping of the schemas for model training, or need access to source schema data which might be unavailable due to privacy concerns. In this paper we present a novel method, named ReMatch, for matching schemas using retrieval-enhanced Large Language Models (LLMs). Our method avoids the need for predefined mapping, any model training, or access to data in the source database. Our experimental results on large real-world schemas demonstrate that ReMatch is an effective matcher. By eliminating the requirement for training data, ReMatch becomes a viable solution for real-world scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2403_01567
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ReMatch: Retrieval Enhanced Schema Matching with LLMs
Sheetrit, Eitam
Brief, Menachem
Mishaeli, Moshik
Elisha, Oren
Databases
Artificial Intelligence
Schema matching is a crucial task in data integration, involving the alignment of a source schema with a target schema to establish correspondence between their elements. This task is challenging due to textual and semantic heterogeneity, as well as differences in schema sizes. Although machine-learning-based solutions have been explored in numerous studies, they often suffer from low accuracy, require manual mapping of the schemas for model training, or need access to source schema data which might be unavailable due to privacy concerns. In this paper we present a novel method, named ReMatch, for matching schemas using retrieval-enhanced Large Language Models (LLMs). Our method avoids the need for predefined mapping, any model training, or access to data in the source database. Our experimental results on large real-world schemas demonstrate that ReMatch is an effective matcher. By eliminating the requirement for training data, ReMatch becomes a viable solution for real-world scenarios.
title ReMatch: Retrieval Enhanced Schema Matching with LLMs
topic Databases
Artificial Intelligence
url https://arxiv.org/abs/2403.01567