Saved in:
Bibliographic Details
Main Author: Kalyan Inturi
Format: Recurso digital
Language:
Published: Zenodo 2026
Online Access:https://doi.org/10.5281/zenodo.18434589
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • <p>Large-scale artificial intelligence (AI) training and inference systems rely on highly available acceleratorinfrastructure, where hardware faults directly reduce effective compute capacity and prolong recovery times. As GPU clusters grow in scale and architectural diversity</p>