Flow with FlorDB: Incremental Context Maintenance for the Machine Learning Lifecycle

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Garcia, Rolando, Kallanagoudar, Pragya, Anand, Chithra, Chasins, Sarah E., Hellerstein, Joseph M., Kerrison, Erin Michelle Turner, Parameswaran, Aditya G.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910700446679040
author Garcia, Rolando
Kallanagoudar, Pragya
Anand, Chithra
Chasins, Sarah E.
Hellerstein, Joseph M.
Kerrison, Erin Michelle Turner
Parameswaran, Aditya G.
author_facet Garcia, Rolando
Kallanagoudar, Pragya
Anand, Chithra
Chasins, Sarah E.
Hellerstein, Joseph M.
Kerrison, Erin Michelle Turner
Parameswaran, Aditya G.
contents In this paper we present techniques to incrementally harvest and query arbitrary metadata from machine learning pipelines, without disrupting agile practices. We center our approach on the developer-favored technique for generating metadata -- log statements -- leveraging the fact that logging creates context. We show how hindsight logging allows such statements to be added and executed post-hoc, without requiring developer foresight. Relational views of incomplete metadata can be queried to dynamically materialize new metadata in bulk and on demand across multiple versions of workflows. This is done in a "metadata later" style, off the critical path of agile development. We realize these ideas in a system called FlorDB and demonstrate how the data context framework covers a range of both ad-hoc metadata as well as special cases treated today by bespoke feature stores and model repositories. Through a usage scenario -- including both ML and human feedback -- we illustrate how the component techniques come together to resolve classic software engineering trade-offs between agility and discipline.
format Preprint
id arxiv_https___arxiv_org_abs_2408_02498
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Flow with FlorDB: Incremental Context Maintenance for the Machine Learning Lifecycle
Garcia, Rolando
Kallanagoudar, Pragya
Anand, Chithra
Chasins, Sarah E.
Hellerstein, Joseph M.
Kerrison, Erin Michelle Turner
Parameswaran, Aditya G.
Databases
Software Engineering
In this paper we present techniques to incrementally harvest and query arbitrary metadata from machine learning pipelines, without disrupting agile practices. We center our approach on the developer-favored technique for generating metadata -- log statements -- leveraging the fact that logging creates context. We show how hindsight logging allows such statements to be added and executed post-hoc, without requiring developer foresight. Relational views of incomplete metadata can be queried to dynamically materialize new metadata in bulk and on demand across multiple versions of workflows. This is done in a "metadata later" style, off the critical path of agile development. We realize these ideas in a system called FlorDB and demonstrate how the data context framework covers a range of both ad-hoc metadata as well as special cases treated today by bespoke feature stores and model repositories. Through a usage scenario -- including both ML and human feedback -- we illustrate how the component techniques come together to resolve classic software engineering trade-offs between agility and discipline.
title Flow with FlorDB: Incremental Context Maintenance for the Machine Learning Lifecycle
topic Databases
Software Engineering
url https://arxiv.org/abs/2408.02498