Towards CXL Resilience to CPU Failures

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Psistakis, Antonis, Ocalan, Burak, Alverti, Chloe, Chaix, Fabien, Alagappan, Ramnatthan, Torrellas, Josep
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914315455430656
author Psistakis, Antonis
Ocalan, Burak
Alverti, Chloe
Chaix, Fabien
Alagappan, Ramnatthan
Torrellas, Josep
author_facet Psistakis, Antonis
Ocalan, Burak
Alverti, Chloe
Chaix, Fabien
Alagappan, Ramnatthan
Torrellas, Josep
contents Compute Express Link (CXL) 3.0 and beyond allows the compute nodes of a cluster to share data with hardware cache coherence and at the granularity of a cache line. This enables shared-memory semantics for distributed computing, but introduces new resilience challenges: a node failure leads to the loss of the dirty data in its caches, corrupting application state. Unfortunately, the CXL specification does not consider processor failures. Moreover, when a component fails, the specification tries to isolate it and continue application execution; there is no attempt to bring the application to a consistent state. To address these limitations, this paper extends the CXL specification to be resilient to node failures, and to correctly recover the application after node failures. We call the system ReCXL. To handle the failure of nodes, ReCXL augments the coherence transaction of a write with messages that propagate the update to a small set of other nodes (i.e., Replicas). Replicas save the update in a hardware Logging Unit. Such replication ensures resilience to node failures. Then, at regular intervals, the Logging Units dump the updates to memory. Recovery involves using the logs in the Logging Units to bring the directory and memory to a correct state. Our evaluation shows that ReCXL enables fault-tolerant execution with only a 30% slowdown over the same platform with no fault-tolerance support.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08271
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards CXL Resilience to CPU Failures
Psistakis, Antonis
Ocalan, Burak
Alverti, Chloe
Chaix, Fabien
Alagappan, Ramnatthan
Torrellas, Josep
Distributed, Parallel, and Cluster Computing
Compute Express Link (CXL) 3.0 and beyond allows the compute nodes of a cluster to share data with hardware cache coherence and at the granularity of a cache line. This enables shared-memory semantics for distributed computing, but introduces new resilience challenges: a node failure leads to the loss of the dirty data in its caches, corrupting application state. Unfortunately, the CXL specification does not consider processor failures. Moreover, when a component fails, the specification tries to isolate it and continue application execution; there is no attempt to bring the application to a consistent state. To address these limitations, this paper extends the CXL specification to be resilient to node failures, and to correctly recover the application after node failures. We call the system ReCXL. To handle the failure of nodes, ReCXL augments the coherence transaction of a write with messages that propagate the update to a small set of other nodes (i.e., Replicas). Replicas save the update in a hardware Logging Unit. Such replication ensures resilience to node failures. Then, at regular intervals, the Logging Units dump the updates to memory. Recovery involves using the logs in the Logging Units to bring the directory and memory to a correct state. Our evaluation shows that ReCXL enables fault-tolerant execution with only a 30% slowdown over the same platform with no fault-tolerance support.
title Towards CXL Resilience to CPU Failures
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2602.08271