BoostER: Leveraging Large Language Models for Enhancing Entity Resolution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Huahang, Li, Shuangyin, Hao, Fei, Zhang, Chen Jason, Song, Yuanfeng, Chen, Lei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917610243751936
author Li, Huahang
Li, Shuangyin
Hao, Fei
Zhang, Chen Jason
Song, Yuanfeng
Chen, Lei
author_facet Li, Huahang
Li, Shuangyin
Hao, Fei
Zhang, Chen Jason
Song, Yuanfeng
Chen, Lei
contents Entity resolution, which involves identifying and merging records that refer to the same real-world entity, is a crucial task in areas like Web data integration. This importance is underscored by the presence of numerous duplicated and multi-version data resources on the Web. However, achieving high-quality entity resolution typically demands significant effort. The advent of Large Language Models (LLMs) like GPT-4 has demonstrated advanced linguistic capabilities, which can be a new paradigm for this task. In this paper, we propose a demonstration system named BoostER that examines the possibility of leveraging LLMs in the entity resolution process, revealing advantages in both easy deployment and low cost. Our approach optimally selects a set of matching questions and poses them to LLMs for verification, then refines the distribution of entity resolution results with the response of LLMs. This offers promising prospects to achieve a high-quality entity resolution result for real-world applications, especially to individuals or small companies without the need for extensive model training or significant financial investment.
format Preprint
id arxiv_https___arxiv_org_abs_2403_06434
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BoostER: Leveraging Large Language Models for Enhancing Entity Resolution
Li, Huahang
Li, Shuangyin
Hao, Fei
Zhang, Chen Jason
Song, Yuanfeng
Chen, Lei
Databases
Entity resolution, which involves identifying and merging records that refer to the same real-world entity, is a crucial task in areas like Web data integration. This importance is underscored by the presence of numerous duplicated and multi-version data resources on the Web. However, achieving high-quality entity resolution typically demands significant effort. The advent of Large Language Models (LLMs) like GPT-4 has demonstrated advanced linguistic capabilities, which can be a new paradigm for this task. In this paper, we propose a demonstration system named BoostER that examines the possibility of leveraging LLMs in the entity resolution process, revealing advantages in both easy deployment and low cost. Our approach optimally selects a set of matching questions and poses them to LLMs for verification, then refines the distribution of entity resolution results with the response of LLMs. This offers promising prospects to achieve a high-quality entity resolution result for real-world applications, especially to individuals or small companies without the need for extensive model training or significant financial investment.
title BoostER: Leveraging Large Language Models for Enhancing Entity Resolution
topic Databases
url https://arxiv.org/abs/2403.06434