Probabilistic Runtime Verification, Evaluation and Risk Assessment of Visual Deep Learning Systems

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Torpmann-Hagen, Birk, Halvorsen, Pål, Riegler, Michael A., Johansen, Dag
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909802849894400
author Torpmann-Hagen, Birk
Halvorsen, Pål
Riegler, Michael A.
Johansen, Dag
author_facet Torpmann-Hagen, Birk
Halvorsen, Pål
Riegler, Michael A.
Johansen, Dag
contents Despite achieving excellent performance on benchmarks, deep neural networks often underperform in real-world deployment due to sensitivity to minor, often imperceptible shifts in input data, known as distributional shifts. These shifts are common in practical scenarios but are rarely accounted for during evaluation, leading to inflated performance metrics. To address this gap, we propose a novel methodology for the verification, evaluation, and risk assessment of deep learning systems. Our approach explicitly models the incidence of distributional shifts at runtime by estimating their probability from outputs of out-of-distribution detectors. We combine these estimates with conditional probabilities of network correctness, structuring them in a binary tree. By traversing this tree, we can compute credible and precise estimates of network accuracy. We assess our approach on five different datasets, with which we simulate deployment conditions characterized by differing frequencies of distributional shift. Our approach consistently outperforms conventional evaluation, with accuracy estimation errors typically ranging between 0.01 and 0.1. We further showcase the potential of our approach on a medical segmentation benchmark, wherein we apply our methods towards risk assessment by associating costs with tree nodes, informing cost-benefit analyses and value-judgments. Ultimately, our approach offers a robust framework for improving the reliability and trustworthiness of deep learning systems, particularly in safety-critical applications, by providing more accurate performance estimates and actionable risk assessments.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19419
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Probabilistic Runtime Verification, Evaluation and Risk Assessment of Visual Deep Learning Systems
Torpmann-Hagen, Birk
Halvorsen, Pål
Riegler, Michael A.
Johansen, Dag
Machine Learning
Artificial Intelligence
Despite achieving excellent performance on benchmarks, deep neural networks often underperform in real-world deployment due to sensitivity to minor, often imperceptible shifts in input data, known as distributional shifts. These shifts are common in practical scenarios but are rarely accounted for during evaluation, leading to inflated performance metrics. To address this gap, we propose a novel methodology for the verification, evaluation, and risk assessment of deep learning systems. Our approach explicitly models the incidence of distributional shifts at runtime by estimating their probability from outputs of out-of-distribution detectors. We combine these estimates with conditional probabilities of network correctness, structuring them in a binary tree. By traversing this tree, we can compute credible and precise estimates of network accuracy. We assess our approach on five different datasets, with which we simulate deployment conditions characterized by differing frequencies of distributional shift. Our approach consistently outperforms conventional evaluation, with accuracy estimation errors typically ranging between 0.01 and 0.1. We further showcase the potential of our approach on a medical segmentation benchmark, wherein we apply our methods towards risk assessment by associating costs with tree nodes, informing cost-benefit analyses and value-judgments. Ultimately, our approach offers a robust framework for improving the reliability and trustworthiness of deep learning systems, particularly in safety-critical applications, by providing more accurate performance estimates and actionable risk assessments.
title Probabilistic Runtime Verification, Evaluation and Risk Assessment of Visual Deep Learning Systems
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.19419