FaaS and Furious: abstractions and differential caching for efficient data pre-processing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tagliabue, Jacopo, Curtin, Ryan, Greco, Ciro
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909386821074944
author Tagliabue, Jacopo
Curtin, Ryan
Greco, Ciro
author_facet Tagliabue, Jacopo
Curtin, Ryan
Greco, Ciro
contents Data pre-processing pipelines are the bread and butter of any successful AI project. We introduce a novel programming model for pipelines in a data lakehouse, allowing users to interact declaratively with assets in object storage. Motivated by real-world industry usage patterns, we exploit these new abstractions with a columnar and differential cache to maximize iteration speed for data scientists, who spent most of their time in pre-processing - adding or removing features, restricting or relaxing time windows, wrangling current or older datasets. We show how the new cache works transparently across programming languages, schemas and time windows, and provide preliminary evidence on its efficiency on standard data workloads.
format Preprint
id arxiv_https___arxiv_org_abs_2411_08203
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FaaS and Furious: abstractions and differential caching for efficient data pre-processing
Tagliabue, Jacopo
Curtin, Ryan
Greco, Ciro
Databases
Data pre-processing pipelines are the bread and butter of any successful AI project. We introduce a novel programming model for pipelines in a data lakehouse, allowing users to interact declaratively with assets in object storage. Motivated by real-world industry usage patterns, we exploit these new abstractions with a columnar and differential cache to maximize iteration speed for data scientists, who spent most of their time in pre-processing - adding or removing features, restricting or relaxing time windows, wrangling current or older datasets. We show how the new cache works transparently across programming languages, schemas and time windows, and provide preliminary evidence on its efficiency on standard data workloads.
title FaaS and Furious: abstractions and differential caching for efficient data pre-processing
topic Databases
url https://arxiv.org/abs/2411.08203