VeML: An End-to-End Machine Learning Lifecycle for Large-scale and High-dimensional Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Le, Van-Duc, Bui, Tien-Cuong, Li, Wen-Syan
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909917902798848
author Le, Van-Duc
Bui, Tien-Cuong
Li, Wen-Syan
author_facet Le, Van-Duc
Bui, Tien-Cuong
Li, Wen-Syan
contents An end-to-end machine learning (ML) lifecycle consists of many iterative processes, from data preparation and ML model design to model training and then deploying the trained model for inference. When building an end-to-end lifecycle for an ML problem, many ML pipelines must be designed and executed that produce a huge number of lifecycle versions. Therefore, this paper introduces VeML, a Version management system dedicated to end-to-end ML Lifecycle. Our system tackles several crucial problems that other systems have not solved. First, we address the high cost of building an ML lifecycle, especially for large-scale and high-dimensional dataset. We solve this problem by proposing to transfer the lifecycle of similar datasets managed in our system to the new training data. We design an algorithm based on the core set to compute similarity for large-scale, high-dimensional data efficiently. Another critical issue is the model accuracy degradation by the difference between training data and testing data during the ML lifetime, which leads to lifecycle rebuild. Our system helps to detect this mismatch without getting labeled data from testing data and rebuild the ML lifecycle for a new data version. To demonstrate our contributions, we conduct experiments on real-world, large-scale datasets of driving images and spatiotemporal sensor data and show promising results.
format Preprint
id arxiv_https___arxiv_org_abs_2304_13037
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle VeML: An End-to-End Machine Learning Lifecycle for Large-scale and High-dimensional Data
Le, Van-Duc
Bui, Tien-Cuong
Li, Wen-Syan
Machine Learning
Databases
Human-Computer Interaction
An end-to-end machine learning (ML) lifecycle consists of many iterative processes, from data preparation and ML model design to model training and then deploying the trained model for inference. When building an end-to-end lifecycle for an ML problem, many ML pipelines must be designed and executed that produce a huge number of lifecycle versions. Therefore, this paper introduces VeML, a Version management system dedicated to end-to-end ML Lifecycle. Our system tackles several crucial problems that other systems have not solved. First, we address the high cost of building an ML lifecycle, especially for large-scale and high-dimensional dataset. We solve this problem by proposing to transfer the lifecycle of similar datasets managed in our system to the new training data. We design an algorithm based on the core set to compute similarity for large-scale, high-dimensional data efficiently. Another critical issue is the model accuracy degradation by the difference between training data and testing data during the ML lifetime, which leads to lifecycle rebuild. Our system helps to detect this mismatch without getting labeled data from testing data and rebuild the ML lifecycle for a new data version. To demonstrate our contributions, we conduct experiments on real-world, large-scale datasets of driving images and spatiotemporal sensor data and show promising results.
title VeML: An End-to-End Machine Learning Lifecycle for Large-scale and High-dimensional Data
topic Machine Learning
Databases
Human-Computer Interaction
url https://arxiv.org/abs/2304.13037