Fast and Adaptive Bulk Loading of Multidimensional Points

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Moti, Moin Hussain, Papadias, Dimitris
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910603932598272
author Moti, Moin Hussain
Papadias, Dimitris
author_facet Moti, Moin Hussain
Papadias, Dimitris
contents Existing methods for bulk loading disk-based multidimensional points involve multiple applications of external sorting. In this paper, we propose techniques that apply linear scan, and are therefore significantly faster. The resulting FMBI Index possesses several desirable properties, including almost full and square nodes with zero overlap, and has excellent query performance. As a second contribution, we develop an adaptive version AMBI, which utilizes the query workload to build a partial index only for parts of the data space that contain query results. Finally, we extend FMBI and AMBI to parallel bulk loading and query processing in distributed systems. An extensive experimental evaluation with real datasets confirms that FMBI and AMBI clearly outperform competitors in terms of combined index construction and query processing cost, sometimes by orders of magnitude.
format Preprint
id arxiv_https___arxiv_org_abs_2409_09447
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fast and Adaptive Bulk Loading of Multidimensional Points
Moti, Moin Hussain
Papadias, Dimitris
Databases
Existing methods for bulk loading disk-based multidimensional points involve multiple applications of external sorting. In this paper, we propose techniques that apply linear scan, and are therefore significantly faster. The resulting FMBI Index possesses several desirable properties, including almost full and square nodes with zero overlap, and has excellent query performance. As a second contribution, we develop an adaptive version AMBI, which utilizes the query workload to build a partial index only for parts of the data space that contain query results. Finally, we extend FMBI and AMBI to parallel bulk loading and query processing in distributed systems. An extensive experimental evaluation with real datasets confirms that FMBI and AMBI clearly outperform competitors in terms of combined index construction and query processing cost, sometimes by orders of magnitude.
title Fast and Adaptive Bulk Loading of Multidimensional Points
topic Databases
url https://arxiv.org/abs/2409.09447