Taming Silent Failures: A Framework for Verifiable AI Reliability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Guan-Yan, Wang, Farn
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914362864697344
author Yang, Guan-Yan
Wang, Farn
author_facet Yang, Guan-Yan
Wang, Farn
contents The integration of Artificial Intelligence (AI) into safety-critical systems introduces a new reliability paradigm: silent failures, where AI produces confident but incorrect outputs that can be dangerous. This paper introduces the Formal Assurance and Monitoring Environment (FAME), a novel framework that confronts this challenge. FAME synergizes the mathematical rigor of offline formal synthesis with the vigilance of online runtime monitoring to create a verifiable safety net around opaque AI components. We demonstrate its efficacy in an autonomous vehicle perception system, where FAME successfully detected 93.5% of critical safety violations that were otherwise silent. By contextualizing our framework within the ISO 26262 and ISO/PAS 8800 standards, we provide reliability engineers with a practical, certifiable pathway for deploying trustworthy AI. FAME represents a crucial shift from accepting probabilistic performance to enforcing provable safety in next-generation systems.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22224
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Taming Silent Failures: A Framework for Verifiable AI Reliability
Yang, Guan-Yan
Wang, Farn
Software Engineering
Artificial Intelligence
Machine Learning
Logic in Computer Science
Systems and Control
The integration of Artificial Intelligence (AI) into safety-critical systems introduces a new reliability paradigm: silent failures, where AI produces confident but incorrect outputs that can be dangerous. This paper introduces the Formal Assurance and Monitoring Environment (FAME), a novel framework that confronts this challenge. FAME synergizes the mathematical rigor of offline formal synthesis with the vigilance of online runtime monitoring to create a verifiable safety net around opaque AI components. We demonstrate its efficacy in an autonomous vehicle perception system, where FAME successfully detected 93.5% of critical safety violations that were otherwise silent. By contextualizing our framework within the ISO 26262 and ISO/PAS 8800 standards, we provide reliability engineers with a practical, certifiable pathway for deploying trustworthy AI. FAME represents a crucial shift from accepting probabilistic performance to enforcing provable safety in next-generation systems.
title Taming Silent Failures: A Framework for Verifiable AI Reliability
topic Software Engineering
Artificial Intelligence
Machine Learning
Logic in Computer Science
Systems and Control
url https://arxiv.org/abs/2510.22224