Should I Hide My Duck in the Lake?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dann, Jonas, Alonso, Gustavo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915980481921024
author Dann, Jonas
Alonso, Gustavo
author_facet Dann, Jonas
Alonso, Gustavo
contents Data lakes spend a significant fraction of query execution time on scanning data from remote, disaggregated storage. Decoding alone accounts for 46% of runtime when running TPC-H directly on Parquet files. To address this bottleneck, we propose a vision for a data processing SmartNIC for the cloud that sits on the network datapath of compute nodes to offload decoding and pushed-down operators, effectively hiding the cost of parsing raw files. Our experimental estimations with DuckDB suggest that by operating directly on pre-filtered data, as delivered by a SmartNIC, we can significantly increase query processing performance and can still match query throughput of traditional setups with smaller, less expensive CPUs.
format Preprint
id arxiv_https___arxiv_org_abs_2602_18775
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Should I Hide My Duck in the Lake?
Dann, Jonas
Alonso, Gustavo
Databases
Data lakes spend a significant fraction of query execution time on scanning data from remote, disaggregated storage. Decoding alone accounts for 46% of runtime when running TPC-H directly on Parquet files. To address this bottleneck, we propose a vision for a data processing SmartNIC for the cloud that sits on the network datapath of compute nodes to offload decoding and pushed-down operators, effectively hiding the cost of parsing raw files. Our experimental estimations with DuckDB suggest that by operating directly on pre-filtered data, as delivered by a SmartNIC, we can significantly increase query processing performance and can still match query throughput of traditional setups with smaller, less expensive CPUs.
title Should I Hide My Duck in the Lake?
topic Databases
url https://arxiv.org/abs/2602.18775