Monitoring Machine Learning Systems: A Multivocal Literature Review

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Naveed, Hira, Barnett, Scott, Arora, Chetan, Grundy, John, Khalajzadeh, Hourieh, Haggag, Omar
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914044211888128
author Naveed, Hira
Barnett, Scott
Arora, Chetan
Grundy, John
Khalajzadeh, Hourieh
Haggag, Omar
author_facet Naveed, Hira
Barnett, Scott
Arora, Chetan
Grundy, John
Khalajzadeh, Hourieh
Haggag, Omar
contents Context: Dynamic production environments make it challenging to maintain reliable machine learning (ML) systems. Runtime issues, such as changes in data patterns or operating contexts, that degrade model performance are a common occurrence in production settings. Monitoring enables early detection and mitigation of these runtime issues, helping maintain users' trust and prevent unwanted consequences for organizations. Aim: This study aims to provide a comprehensive overview of the ML monitoring literature. Method: We conducted a multivocal literature review (MLR) following the well established guidelines by Garousi to investigate various aspects of ML monitoring approaches in 136 papers. Results: We analyzed selected studies based on four key areas: (1) the motivations, goals, and context; (2) the monitored aspects, specific techniques, metrics, and tools; (3) the contributions and benefits; and (4) the current limitations. We also discuss several insights found in the studies, their implications, and recommendations for future research and practice. Conclusion: Our MLR identifies and summarizes ML monitoring practices and gaps, emphasizing similarities and disconnects between formal and gray literature. Our study is valuable for both academics and practitioners, as it helps select appropriate solutions, highlights limitations in current approaches, and provides future directions for research and tool development.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14294
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Monitoring Machine Learning Systems: A Multivocal Literature Review
Naveed, Hira
Barnett, Scott
Arora, Chetan
Grundy, John
Khalajzadeh, Hourieh
Haggag, Omar
Software Engineering
Machine Learning
Context: Dynamic production environments make it challenging to maintain reliable machine learning (ML) systems. Runtime issues, such as changes in data patterns or operating contexts, that degrade model performance are a common occurrence in production settings. Monitoring enables early detection and mitigation of these runtime issues, helping maintain users' trust and prevent unwanted consequences for organizations. Aim: This study aims to provide a comprehensive overview of the ML monitoring literature. Method: We conducted a multivocal literature review (MLR) following the well established guidelines by Garousi to investigate various aspects of ML monitoring approaches in 136 papers. Results: We analyzed selected studies based on four key areas: (1) the motivations, goals, and context; (2) the monitored aspects, specific techniques, metrics, and tools; (3) the contributions and benefits; and (4) the current limitations. We also discuss several insights found in the studies, their implications, and recommendations for future research and practice. Conclusion: Our MLR identifies and summarizes ML monitoring practices and gaps, emphasizing similarities and disconnects between formal and gray literature. Our study is valuable for both academics and practitioners, as it helps select appropriate solutions, highlights limitations in current approaches, and provides future directions for research and tool development.
title Monitoring Machine Learning Systems: A Multivocal Literature Review
topic Software Engineering
Machine Learning
url https://arxiv.org/abs/2509.14294