Trustworthy AI in the Agentic Lakehouse: from Concurrency to Governance

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tagliabue, Jacopo, Bianchi, Federico, Greco, Ciro
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918211774054400
author Tagliabue, Jacopo
Bianchi, Federico
Greco, Ciro
author_facet Tagliabue, Jacopo
Bianchi, Federico
Greco, Ciro
contents Even as AI capabilities improve, most enterprises do not consider agents trustworthy enough to work on production data. In this paper, we argue that the path to trustworthy agentic workflows begins with solving the infrastructure problem first: traditional lakehouses are not suited for agent access patterns, but if we design one around transactions, governance follows. In particular, we draw an operational analogy to MVCC in databases and show why a direct transplant fails in a decoupled, multi-language setting. We then propose an agent-first design, Bauplan, that reimplements data and compute isolation in the lakehouse. We conclude by sharing a reference implementation of a self-healing pipeline in Bauplan, which seamlessly couples agent reasoning with all the desired guarantees for correctness and trust.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16402
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Trustworthy AI in the Agentic Lakehouse: from Concurrency to Governance
Tagliabue, Jacopo
Bianchi, Federico
Greco, Ciro
Artificial Intelligence
Databases
Even as AI capabilities improve, most enterprises do not consider agents trustworthy enough to work on production data. In this paper, we argue that the path to trustworthy agentic workflows begins with solving the infrastructure problem first: traditional lakehouses are not suited for agent access patterns, but if we design one around transactions, governance follows. In particular, we draw an operational analogy to MVCC in databases and show why a direct transplant fails in a decoupled, multi-language setting. We then propose an agent-first design, Bauplan, that reimplements data and compute isolation in the lakehouse. We conclude by sharing a reference implementation of a self-healing pipeline in Bauplan, which seamlessly couples agent reasoning with all the desired guarantees for correctness and trust.
title Trustworthy AI in the Agentic Lakehouse: from Concurrency to Governance
topic Artificial Intelligence
Databases
url https://arxiv.org/abs/2511.16402