Mind the Data Gap: Bridging LLMs to Enterprise Data Integration

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kayali, Moe, Wenz, Fabian, Tatbul, Nesime, Demiralp, Çağatay
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916545266974720
author Kayali, Moe
Wenz, Fabian
Tatbul, Nesime
Demiralp, Çağatay
author_facet Kayali, Moe
Wenz, Fabian
Tatbul, Nesime
Demiralp, Çağatay
contents Leading large language models (LLMs) are trained on public data. However, most of the world's data is dark data that is not publicly accessible, mainly in the form of private organizational or enterprise data. We show that the performance of methods based on LLMs seriously degrades when tested on real-world enterprise datasets. Current benchmarks, based on public data, overestimate the performance of LLMs. We release a new benchmark dataset, the GOBY Benchmark, to advance discovery in enterprise data integration. Based on our experience with this enterprise benchmark, we propose techniques to uplift the performance of LLMs on enterprise data, including (1) hierarchical annotation, (2) runtime class-learning, and (3) ontology synthesis. We show that, once these techniques are deployed, the performance on enterprise data becomes on par with that of public data. The Goby benchmark can be obtained at https://goby-benchmark.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20331
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mind the Data Gap: Bridging LLMs to Enterprise Data Integration
Kayali, Moe
Wenz, Fabian
Tatbul, Nesime
Demiralp, Çağatay
Databases
Artificial Intelligence
Machine Learning
Leading large language models (LLMs) are trained on public data. However, most of the world's data is dark data that is not publicly accessible, mainly in the form of private organizational or enterprise data. We show that the performance of methods based on LLMs seriously degrades when tested on real-world enterprise datasets. Current benchmarks, based on public data, overestimate the performance of LLMs. We release a new benchmark dataset, the GOBY Benchmark, to advance discovery in enterprise data integration. Based on our experience with this enterprise benchmark, we propose techniques to uplift the performance of LLMs on enterprise data, including (1) hierarchical annotation, (2) runtime class-learning, and (3) ontology synthesis. We show that, once these techniques are deployed, the performance on enterprise data becomes on par with that of public data. The Goby benchmark can be obtained at https://goby-benchmark.github.io/.
title Mind the Data Gap: Bridging LLMs to Enterprise Data Integration
topic Databases
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.20331