TensorBank: Tensor Lakehouse for Foundation Model Training

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kienzler, Romeo, Tizzei, Leonardo Pondian, Blumenstiel, Benedikt, Nagy, Zoltan Arnold, Mukkavilli, S. Karthik, Schmude, Johannes, Freitag, Marcus, Behrendt, Michael, Civitarese, Daniel Salles, Simumba, Naomi, Kimura, Daiki, Hamann, Hendrik
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911806205722624
author Kienzler, Romeo
Tizzei, Leonardo Pondian
Blumenstiel, Benedikt
Nagy, Zoltan Arnold
Mukkavilli, S. Karthik
Schmude, Johannes
Freitag, Marcus
Behrendt, Michael
Civitarese, Daniel Salles
Simumba, Naomi
Kimura, Daiki
Hamann, Hendrik
author_facet Kienzler, Romeo
Tizzei, Leonardo Pondian
Blumenstiel, Benedikt
Nagy, Zoltan Arnold
Mukkavilli, S. Karthik
Schmude, Johannes
Freitag, Marcus
Behrendt, Michael
Civitarese, Daniel Salles
Simumba, Naomi
Kimura, Daiki
Hamann, Hendrik
contents Storing and streaming high dimensional data for foundation model training became a critical requirement with the rise of foundation models beyond natural language. In this paper we introduce TensorBank, a petabyte scale tensor lakehouse capable of streaming tensors from Cloud Object Store (COS) to GPU memory at wire speed based on complex relational queries. We use Hierarchical Statistical Indices (HSI) for query acceleration. Our architecture allows to directly address tensors on block level using HTTP range reads. Once in GPU memory, data can be transformed using PyTorch transforms. We provide a generic PyTorch dataset type with a corresponding dataset factory translating relational queries and requested transformations as an instance. By making use of the HSI, irrelevant blocks can be skipped without reading them as those indices contain statistics on their content at different hierarchical resolution levels. This is an opinionated architecture powered by open standards and making heavy use of open-source technology. Although, hardened for production use using geospatial-temporal data, this architecture generalizes to other use case like computer vision, computational neuroscience, biological sequence analysis and more.
format Preprint
id arxiv_https___arxiv_org_abs_2309_02094
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle TensorBank: Tensor Lakehouse for Foundation Model Training
Kienzler, Romeo
Tizzei, Leonardo Pondian
Blumenstiel, Benedikt
Nagy, Zoltan Arnold
Mukkavilli, S. Karthik
Schmude, Johannes
Freitag, Marcus
Behrendt, Michael
Civitarese, Daniel Salles
Simumba, Naomi
Kimura, Daiki
Hamann, Hendrik
Machine Learning
Artificial Intelligence
Databases
Information Retrieval
Storing and streaming high dimensional data for foundation model training became a critical requirement with the rise of foundation models beyond natural language. In this paper we introduce TensorBank, a petabyte scale tensor lakehouse capable of streaming tensors from Cloud Object Store (COS) to GPU memory at wire speed based on complex relational queries. We use Hierarchical Statistical Indices (HSI) for query acceleration. Our architecture allows to directly address tensors on block level using HTTP range reads. Once in GPU memory, data can be transformed using PyTorch transforms. We provide a generic PyTorch dataset type with a corresponding dataset factory translating relational queries and requested transformations as an instance. By making use of the HSI, irrelevant blocks can be skipped without reading them as those indices contain statistics on their content at different hierarchical resolution levels. This is an opinionated architecture powered by open standards and making heavy use of open-source technology. Although, hardened for production use using geospatial-temporal data, this architecture generalizes to other use case like computer vision, computational neuroscience, biological sequence analysis and more.
title TensorBank: Tensor Lakehouse for Foundation Model Training
topic Machine Learning
Artificial Intelligence
Databases
Information Retrieval
url https://arxiv.org/abs/2309.02094