Instance-Optimized String Fingerprints

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Stoian, Mihail, Thürauf, Johannes, Zimmerer, Andreas, van Renen, Alexander, Kipf, Andreas
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912481826308096
author Stoian, Mihail
Thürauf, Johannes
Zimmerer, Andreas
van Renen, Alexander
Kipf, Andreas
author_facet Stoian, Mihail
Thürauf, Johannes
Zimmerer, Andreas
van Renen, Alexander
Kipf, Andreas
contents Recent research found that cloud data warehouses are text-heavy. However, their capabilities for efficiently processing string columns remain limited, relying primarily on techniques like dictionary encoding and prefix-based partition pruning. In recent work, we introduced string fingerprints - a lightweight secondary index structure designed to approximate LIKE predicates, albeit with false positives. This approach is particularly compelling for columnar query engines, where fingerprints can help reduce both compute and I/O overhead. We show that string fingerprints can be optimized for specific workloads using mixed-integer optimization, and that they can generalize to unseen table predicates. On an IMDb column evaluated in DuckDB v1.3, this yields table-scan speedups of up to 1.36$\times$.
format Preprint
id arxiv_https___arxiv_org_abs_2507_10391
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Instance-Optimized String Fingerprints
Stoian, Mihail
Thürauf, Johannes
Zimmerer, Andreas
van Renen, Alexander
Kipf, Andreas
Databases
Recent research found that cloud data warehouses are text-heavy. However, their capabilities for efficiently processing string columns remain limited, relying primarily on techniques like dictionary encoding and prefix-based partition pruning. In recent work, we introduced string fingerprints - a lightweight secondary index structure designed to approximate LIKE predicates, albeit with false positives. This approach is particularly compelling for columnar query engines, where fingerprints can help reduce both compute and I/O overhead. We show that string fingerprints can be optimized for specific workloads using mixed-integer optimization, and that they can generalize to unseen table predicates. On an IMDb column evaluated in DuckDB v1.3, this yields table-scan speedups of up to 1.36$\times$.
title Instance-Optimized String Fingerprints
topic Databases
url https://arxiv.org/abs/2507.10391