SiMa: Effective and Efficient Matching Across Data Silos Using Graph Neural Networks

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Koutras, Christos, Hai, Rihan, Psarakis, Kyriakos, Fragkoulis, Marios, Katsifodimos, Asterios
Formato: Preprint
Publicado: 2022
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916143991619584
author Koutras, Christos
Hai, Rihan
Psarakis, Kyriakos
Fragkoulis, Marios
Katsifodimos, Asterios
author_facet Koutras, Christos
Hai, Rihan
Psarakis, Kyriakos
Fragkoulis, Marios
Katsifodimos, Asterios
contents How can we leverage existing column relationships within silos, to predict similar ones across silos? Can we do this efficiently and effectively? Existing matching approaches do not exploit prior knowledge, relying on prohibitively expensive similarity computations. In this paper we present the first technique for matching columns across data silos, called SiMa, which leverages Graph Neural Networks (GNNs) to learn from existing column relationships within data silos, and dataset-specific profiles. The main novelty of SiMa is its ability to be trained incrementally on column relationships within each silo individually, without requiring the consolidation of all datasets in a single place. Our experiments show that SiMa is more effective than the - otherwise inapplicable to the setting of silos - state-of-the-art matching methods, while requiring orders of magnitude less computational resources. Moreover, we demonstrate that SiMa considerably outperforms other state-of-the-art column representation learning methods.
format Preprint
id arxiv_https___arxiv_org_abs_2206_12733
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle SiMa: Effective and Efficient Matching Across Data Silos Using Graph Neural Networks
Koutras, Christos
Hai, Rihan
Psarakis, Kyriakos
Fragkoulis, Marios
Katsifodimos, Asterios
Databases
How can we leverage existing column relationships within silos, to predict similar ones across silos? Can we do this efficiently and effectively? Existing matching approaches do not exploit prior knowledge, relying on prohibitively expensive similarity computations. In this paper we present the first technique for matching columns across data silos, called SiMa, which leverages Graph Neural Networks (GNNs) to learn from existing column relationships within data silos, and dataset-specific profiles. The main novelty of SiMa is its ability to be trained incrementally on column relationships within each silo individually, without requiring the consolidation of all datasets in a single place. Our experiments show that SiMa is more effective than the - otherwise inapplicable to the setting of silos - state-of-the-art matching methods, while requiring orders of magnitude less computational resources. Moreover, we demonstrate that SiMa considerably outperforms other state-of-the-art column representation learning methods.
title SiMa: Effective and Efficient Matching Across Data Silos Using Graph Neural Networks
topic Databases
url https://arxiv.org/abs/2206.12733