Automating the Data Science Lifecycle: CI/CD for Machine Learning Deployment

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Suddala, Swathi
Format: Recurso digital
Sprache:Englisch
Veröffentlicht: Zenodo 2022
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866902265579700224
author Suddala, Swathi
author_facet Suddala, Swathi
contents <p>The incorporation of Continuous Integration (CI) and Continuous Deployment (CD) into the machine <br>learning (ML) lifecycle is essential for facilitating the effective transition of models from the <br>development phase to production. Unlike conventional software, ML workflows face distinct challenges <br>such as data versioning, model drift, hyperparameter optimization, and limitations in computational <br>resources. This paper explores optimal practices for automating the data science lifecycle through CI/CD <br>methodologies, focusing on critical elements like automated data validation, model retraining pipelines, <br>and deployment orchestration. We analyze the significance of infrastructure-as-code, Docker <br>containerization, model registries, and monitor frameworks in enhancing ML operations. Additionally, <br>we propose a robust framework that ensures reproducibility, scalability, and reliability in the deployment <br>of ML models. The study also highlights sophisticated CI/CD strategies tailored for machine learning, <br>emphasizing the vital role of MLOps practices in maintaining model integrity within ever-evolving <br>production settings.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_15328138
institution Zenodo
language eng
publishDate 2022
publisher Zenodo
record_format zenodo
spellingShingle Automating the Data Science Lifecycle: CI/CD for Machine Learning Deployment
Suddala, Swathi
Data Science
Data analysis
Machine Learning
data versioning
<p>The incorporation of Continuous Integration (CI) and Continuous Deployment (CD) into the machine <br>learning (ML) lifecycle is essential for facilitating the effective transition of models from the <br>development phase to production. Unlike conventional software, ML workflows face distinct challenges <br>such as data versioning, model drift, hyperparameter optimization, and limitations in computational <br>resources. This paper explores optimal practices for automating the data science lifecycle through CI/CD <br>methodologies, focusing on critical elements like automated data validation, model retraining pipelines, <br>and deployment orchestration. We analyze the significance of infrastructure-as-code, Docker <br>containerization, model registries, and monitor frameworks in enhancing ML operations. Additionally, <br>we propose a robust framework that ensures reproducibility, scalability, and reliability in the deployment <br>of ML models. The study also highlights sophisticated CI/CD strategies tailored for machine learning, <br>emphasizing the vital role of MLOps practices in maintaining model integrity within ever-evolving <br>production settings.</p>
title Automating the Data Science Lifecycle: CI/CD for Machine Learning Deployment
topic Data Science
Data analysis
Machine Learning
data versioning
url https://doi.org/10.5281/zenodo.15328138