Hamava: Fault-tolerant Reconfigurable Geo-Replication on Heterogeneous Clusters
Fuente:
arXiv
Saved in:
| Main Authors: | Mane, Tejas, Li, Xiao, Sadoghi, Mohammad, Lesani, Mohsen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reconfigurable Heterogeneous Quorum Systems
by: Li, Xiao, et al.
Published: (2023)
by: Li, Xiao, et al.
Published: (2023)
SafarDB: FPGA-Accelerated Distributed Transactions via Replicated Data Types
by: Saberlatibari, Javad, et al.
Published: (2026)
by: Saberlatibari, Javad, et al.
Published: (2026)
Fault-tolerant Consensus in Anonymous Dynamic Network
by: Zhang, Qinzi, et al.
Published: (2024)
by: Zhang, Qinzi, et al.
Published: (2024)
Rubick: Exploiting Job Reconfigurability for Deep Learning Cluster Scheduling
by: Zhang, Xinyi, et al.
Published: (2024)
by: Zhang, Xinyi, et al.
Published: (2024)
Fault-tolerant Reduce and Allreduce operations based on correction
by: Kuettler, Martin, et al.
Published: (2026)
by: Kuettler, Martin, et al.
Published: (2026)
The Power of Abstract MAC Layer: A Fault-tolerance Perspective
by: Zhang, Qinzi, et al.
Published: (2024)
by: Zhang, Qinzi, et al.
Published: (2024)
Resilient Packet Forwarding: A Reinforcement Learning Approach to Routing in Gaussian Interconnected Networks with Clustered Faults
by: Charrwi, Mohammad Walid, et al.
Published: (2025)
by: Charrwi, Mohammad Walid, et al.
Published: (2025)
Did we miss P In CAP? Partial Progress Conjecture under Asynchrony
by: Chen, Junchao, et al.
Published: (2024)
by: Chen, Junchao, et al.
Published: (2024)
Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed Clusters
by: Strati, Foteini, et al.
Published: (2025)
by: Strati, Foteini, et al.
Published: (2025)
FASTEN: Towards a FAult-tolerant and STorage EfficieNt Cloud: Balancing Between Replication and Deduplication
by: Ahmed, Sabbir, et al.
Published: (2023)
by: Ahmed, Sabbir, et al.
Published: (2023)
Energy-Aware Scheduling Strategies for Partially-Replicable Task Chains on Heterogeneous Processors
by: Idouar, Yacine, et al.
Published: (2025)
by: Idouar, Yacine, et al.
Published: (2025)
Fides: Secure and Scalable Asynchronous DAG Consensus via Trusted Components
by: Xie, Shaokang, et al.
Published: (2025)
by: Xie, Shaokang, et al.
Published: (2025)
A Logic for Repair and State Recovery in Byzantine Fault-tolerant Multi-agent Systems
by: van Ditmarsch, Hans, et al.
Published: (2024)
by: van Ditmarsch, Hans, et al.
Published: (2024)
DéjàVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving
by: Strati, Foteini, et al.
Published: (2024)
by: Strati, Foteini, et al.
Published: (2024)
Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters
by: Chang, Zihan, et al.
Published: (2024)
by: Chang, Zihan, et al.
Published: (2024)
Spider: A BFT Architecture for Geo-Replicated Cloud Services
by: Eischer, Michael, et al.
Published: (2024)
by: Eischer, Michael, et al.
Published: (2024)
Fast and Interactive Byzantine Fault-tolerant Web Services via Session-Based Consensus Decoupling
by: Akmal, Ahmad Zaki, et al.
Published: (2025)
by: Akmal, Ahmad Zaki, et al.
Published: (2025)
HYDRA: Breaking the Global Ordering Barrier in Multi-BFT Consensus
by: Lyu, Hanzheng, et al.
Published: (2025)
by: Lyu, Hanzheng, et al.
Published: (2025)
Benchmarking Different Application Types across Heterogeneous Cloud Compute Services
by: Duggi, Nivedhitha, et al.
Published: (2025)
by: Duggi, Nivedhitha, et al.
Published: (2025)
Dalek: An Unconventional and Energy-Aware Heterogeneous Cluster
by: Cassagne, Adrien, et al.
Published: (2025)
by: Cassagne, Adrien, et al.
Published: (2025)
HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters
by: Liang, Antian, et al.
Published: (2025)
by: Liang, Antian, et al.
Published: (2025)
LaissezCloud: Continuous Resource Renegotiation for the Public Cloud
by: Harith, Tejas, et al.
Published: (2026)
by: Harith, Tejas, et al.
Published: (2026)
It's the People, Not the Placement: Rethinking Allocations in Post-Moore Clouds
by: Harith, Tejas, et al.
Published: (2025)
by: Harith, Tejas, et al.
Published: (2025)
VBFT: Veloce Byzantine Fault Tolerant Consensus for Blockchains
by: Jalalzai, Mohammad M., et al.
Published: (2023)
by: Jalalzai, Mohammad M., et al.
Published: (2023)
Optimal Resource Efficiency with Fairness in Heterogeneous GPU Clusters
by: Mo, Zizhao, et al.
Published: (2024)
by: Mo, Zizhao, et al.
Published: (2024)
Training DNN Models over Heterogeneous Clusters with Optimal Performance
by: Nie, Chengyi, et al.
Published: (2024)
by: Nie, Chengyi, et al.
Published: (2024)
Cephalo: Harnessing Heterogeneous GPU Clusters for Training Transformer Models
by: Guo, Runsheng Benson, et al.
Published: (2024)
by: Guo, Runsheng Benson, et al.
Published: (2024)
Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters
by: Guo, Runsheng Benson, et al.
Published: (2025)
by: Guo, Runsheng Benson, et al.
Published: (2025)
Orthrus: Accelerating Multi-BFT Consensus through Concurrent Partial Ordering of Transactions (Extended Version)
by: Lyu, Hanzheng, et al.
Published: (2024)
by: Lyu, Hanzheng, et al.
Published: (2024)
Selection Guidelines for Geo-Replicated SMR Protocols: A Communication Pattern-based Latency Modeling Approach
by: Shiozaki, Kohya, et al.
Published: (2024)
by: Shiozaki, Kohya, et al.
Published: (2024)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
by: Zhang, WenZheng, et al.
Published: (2024)
by: Zhang, WenZheng, et al.
Published: (2024)
Augur: Pre-Execution Energy Prediction for Workflow Tasks in Heterogeneous Clusters
by: West, Kathleen, et al.
Published: (2026)
by: West, Kathleen, et al.
Published: (2026)
Undo and Redo Support for Replicated Registers
by: Stewen, Leo, et al.
Published: (2024)
by: Stewen, Leo, et al.
Published: (2024)
Reliable Replication Protocols on SmartNICs
by: Katebzadeh, M. R. Siavash, et al.
Published: (2025)
by: Katebzadeh, M. R. Siavash, et al.
Published: (2025)
HAP: SPMD DNN Training on Heterogeneous GPU Clusters with Automated Program Synthesis
by: Zhang, Shiwei, et al.
Published: (2024)
by: Zhang, Shiwei, et al.
Published: (2024)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
by: Mo, Zizhao, et al.
Published: (2025)
by: Mo, Zizhao, et al.
Published: (2025)
Building State Machine Replication Using Practical Network Synchrony
by: Wan, Yiliang, et al.
Published: (2025)
by: Wan, Yiliang, et al.
Published: (2025)
Self-Stabilizing Replicated State Machine Coping with Byzantine and Recurring Transient Faults
by: Dolev, Shlomi, et al.
Published: (2025)
by: Dolev, Shlomi, et al.
Published: (2025)
Revisiting Speculative Leaderless Protocols for Low-Latency BFT Replication
by: Qian, Daniel, et al.
Published: (2026)
by: Qian, Daniel, et al.
Published: (2026)
Optimizing View Change for Byzantine Fault Tolerance in Parallel Consensus
by: Xie, Yifei, et al.
Published: (2026)
by: Xie, Yifei, et al.
Published: (2026)
Similar Items
-
Reconfigurable Heterogeneous Quorum Systems
by: Li, Xiao, et al.
Published: (2023) -
SafarDB: FPGA-Accelerated Distributed Transactions via Replicated Data Types
by: Saberlatibari, Javad, et al.
Published: (2026) -
Fault-tolerant Consensus in Anonymous Dynamic Network
by: Zhang, Qinzi, et al.
Published: (2024) -
Rubick: Exploiting Job Reconfigurability for Deep Learning Cluster Scheduling
by: Zhang, Xinyi, et al.
Published: (2024) -
Fault-tolerant Reduce and Allreduce operations based on correction
by: Kuettler, Martin, et al.
Published: (2026)