Saved in:
| Main Author: | |
|---|---|
| Format: | Recurso digital |
| Language: | |
| Published: |
Zenodo
2026
|
| Online Access: | https://doi.org/10.5281/zenodo.18434589 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Table of Contents:
- <p>Large-scale artificial intelligence (AI) training and inference systems rely on highly available acceleratorinfrastructure, where hardware faults directly reduce effective compute capacity and prolong recovery times. As GPU clusters grow in scale and architectural diversity</p>