Varuna: Enabling Failure-Type Aware RDMA Failover
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xiaoyang, Li, Yongkun, Yao, Lulu, Wei, Guoli, Yang, Longcheng, Xu, Yinlong, Kong, Weiqing, Wang, Weiguang, Dong, Peng, Liu, Bingyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Uber's Failover Architecture: Reconciling Reliability and Efficiency in Hyperscale Microservice Infrastructure
von: Bansal, Mayank, et al.
Veröffentlicht: (2026)
von: Bansal, Mayank, et al.
Veröffentlicht: (2026)
Reimagining RDMA Through the Lens of ML
von: Warraich, Ertza, et al.
Veröffentlicht: (2025)
von: Warraich, Ertza, et al.
Veröffentlicht: (2025)
On the Resilience of Fast Failover Routing Against Dynamic Link Failures
von: Dai, Wenkai, et al.
Veröffentlicht: (2024)
von: Dai, Wenkai, et al.
Veröffentlicht: (2024)
FedRDMA: Communication-Efficient Cross-Silo Federated LLM via Chunked RDMA Transmission
von: Zhang, Zeling, et al.
Veröffentlicht: (2024)
von: Zhang, Zeling, et al.
Veröffentlicht: (2024)
Closing the HPC-Cloud Convergence Gap: Multi-Tenant Slingshot RDMA for Kubernetes
von: Friese, Philipp A., et al.
Veröffentlicht: (2025)
von: Friese, Philipp A., et al.
Veröffentlicht: (2025)
OptiNIC: A Resilient and Tail-Optimal RDMA NIC for Distributed ML Workloads
von: Warraich, Ertza, et al.
Veröffentlicht: (2025)
von: Warraich, Ertza, et al.
Veröffentlicht: (2025)
Palladium: A DPU-enabled Multi-Tenant Serverless Cloud over Zero-copy Multi-node RDMA Fabrics
von: Qi, Shixiong, et al.
Veröffentlicht: (2025)
von: Qi, Shixiong, et al.
Veröffentlicht: (2025)
FODT: Fast, Online, Distributed and Temporary Failure Recovery Approach for MEC
von: Yuan, Xin, et al.
Veröffentlicht: (2023)
von: Yuan, Xin, et al.
Veröffentlicht: (2023)
Study of Workload Interference with Intelligent Routing on Dragonfly
von: Kang, Yao, et al.
Veröffentlicht: (2024)
von: Kang, Yao, et al.
Veröffentlicht: (2024)
Q-adaptive: A Multi-Agent Reinforcement Learning Based Routing on Dragonfly Network
von: Kang, Yao, et al.
Veröffentlicht: (2024)
von: Kang, Yao, et al.
Veröffentlicht: (2024)
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
von: Du, Chengze, et al.
Veröffentlicht: (2025)
von: Du, Chengze, et al.
Veröffentlicht: (2025)
Enabling Scalability in Asynchronous and Bidirectional Communication in LPWAN
von: Rahman, Mahbubur
Veröffentlicht: (2025)
von: Rahman, Mahbubur
Veröffentlicht: (2025)
Enabling Blockchain Interoperability Through Network Discovery Services
von: Hassan, Khalid, et al.
Veröffentlicht: (2025)
von: Hassan, Khalid, et al.
Veröffentlicht: (2025)
Meili: Enabling SmartNIC as a Service in the Cloud
von: Su, Qiang, et al.
Veröffentlicht: (2023)
von: Su, Qiang, et al.
Veröffentlicht: (2023)
Relay-Based Synchronization of Replicated Data Types in Opportunistic Networks
von: Guidec, Frédéric, et al.
Veröffentlicht: (2026)
von: Guidec, Frédéric, et al.
Veröffentlicht: (2026)
A Lightweight and Scalable Design of Segment Routing in Broadband LEO Constellations Using Landmark-Based Skeleton Graphs
von: Hu, Menglan, et al.
Veröffentlicht: (2024)
von: Hu, Menglan, et al.
Veröffentlicht: (2024)
Legible Consensus: Topology-Aware Quorum Geometry for Asymmetric Networks
von: Mason, Tony
Veröffentlicht: (2026)
von: Mason, Tony
Veröffentlicht: (2026)
Contention-Aware Microservice Deployment in Collaborative Mobile Edge Networks
von: Ge, Xinlei, et al.
Veröffentlicht: (2024)
von: Ge, Xinlei, et al.
Veröffentlicht: (2024)
An Auction-Based Mechanism for Optimal Task Allocation and Resource Aware Containerization
von: kumar, Ramakant
Veröffentlicht: (2026)
von: kumar, Ramakant
Veröffentlicht: (2026)
Carbon-Aware Temporal Data Transfer Scheduling Across Cloud Datacenters
von: Rodrigues, Elvis, et al.
Veröffentlicht: (2025)
von: Rodrigues, Elvis, et al.
Veröffentlicht: (2025)
Resource Allocation Driven by Large Models in Future Semantic-Aware Networks
von: Zhang, Haijun, et al.
Veröffentlicht: (2025)
von: Zhang, Haijun, et al.
Veröffentlicht: (2025)
SDNator is Not Another SDN Controller: Enabling Extensible Data-Driven Control in Cyber-Physical Systems
von: Lin, Y., et al.
Veröffentlicht: (2026)
von: Lin, Y., et al.
Veröffentlicht: (2026)
QoS-Aware Load Balancing in the Computing Continuum via Multi-Player Bandits
von: Čilić, Ivan, et al.
Veröffentlicht: (2025)
von: Čilić, Ivan, et al.
Veröffentlicht: (2025)
Risk-Aware and Stable Edge Server Selection Under Network Latency SLOs
von: Liyanage, Mohan, et al.
Veröffentlicht: (2026)
von: Liyanage, Mohan, et al.
Veröffentlicht: (2026)
An Online Fragmentation-Aware GPU Scheduler for Multi-Tenant MIG-based Clouds
von: Zambianco, Marco, et al.
Veröffentlicht: (2025)
von: Zambianco, Marco, et al.
Veröffentlicht: (2025)
A Task Decomposition and Planning Framework for Efficient LLM Inference in AI-Enabled WiFi-Offload Networks
von: Han, Mingqi, et al.
Veröffentlicht: (2026)
von: Han, Mingqi, et al.
Veröffentlicht: (2026)
Semantic-Aware LLM Orchestration for Proactive Resource Management in Predictive Digital Twin Vehicular Networks
von: Ahmadpanah, Seyed Hossein
Veröffentlicht: (2025)
von: Ahmadpanah, Seyed Hossein
Veröffentlicht: (2025)
Dynamic Edge Server Selection in Time-Varying Environments: A Reliability-Aware Predictive Approach
von: Burbano, Jaime Sebastian, et al.
Veröffentlicht: (2025)
von: Burbano, Jaime Sebastian, et al.
Veröffentlicht: (2025)
RouterWise: Joint Resource Allocation and Routing for Latency-Aware Multi-Model LLM Serving
von: Kasnavieh, Hossein Hosseini, et al.
Veröffentlicht: (2026)
von: Kasnavieh, Hossein Hosseini, et al.
Veröffentlicht: (2026)
A Hybrid Approach to Monitor Context Parameters for Optimising Caching for Context-Aware IoT Applications
von: Manchanda, Ashish, et al.
Veröffentlicht: (2024)
von: Manchanda, Ashish, et al.
Veröffentlicht: (2024)
D-LoRa: a Distributed Parameter Adaptation Scheme for LoRa Network
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications
von: Wang, Xi, et al.
Veröffentlicht: (2025)
von: Wang, Xi, et al.
Veröffentlicht: (2025)
PSMOA: Policy Support Multi-Objective Optimization Algorithm for Decentralized Data Replication
von: Wang, Xi, et al.
Veröffentlicht: (2025)
von: Wang, Xi, et al.
Veröffentlicht: (2025)
Don't Let a Few Network Failures Slow the Entire AllReduce
von: Chen, Peiqing, et al.
Veröffentlicht: (2026)
von: Chen, Peiqing, et al.
Veröffentlicht: (2026)
MOFCO: Mobility- and Migration-Aware Task Offloading in Three-Layer Fog Computing Environments
von: Mahdizadeh, Soheil, et al.
Veröffentlicht: (2025)
von: Mahdizadeh, Soheil, et al.
Veröffentlicht: (2025)
From Skew to Symmetry: Node-Interconnect Multi-Path Balancing with Execution-time Planning for Modern GPU Clusters
von: Yao, Jinghan, et al.
Veröffentlicht: (2026)
von: Yao, Jinghan, et al.
Veröffentlicht: (2026)
CRAFT: Latency and Cost-Aware Genetic-Based Framework for Node Placement in Edge-Fog Environments
von: Mahdizadeh, Soheil, et al.
Veröffentlicht: (2025)
von: Mahdizadeh, Soheil, et al.
Veröffentlicht: (2025)
Design and Operation of Shared Machine Learning Clusters on Campus
von: Xu, Kaiqiang, et al.
Veröffentlicht: (2021)
von: Xu, Kaiqiang, et al.
Veröffentlicht: (2021)
Fast Multichannel Topology Discovery in Cognitive Radio Networks
von: Wang, Yung-Li, et al.
Veröffentlicht: (2025)
von: Wang, Yung-Li, et al.
Veröffentlicht: (2025)
Multi-Source Coflow Scheduling in Collaborative Edge Computing with Multihop Network
von: Sahni, Yuvraj, et al.
Veröffentlicht: (2024)
von: Sahni, Yuvraj, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Uber's Failover Architecture: Reconciling Reliability and Efficiency in Hyperscale Microservice Infrastructure
von: Bansal, Mayank, et al.
Veröffentlicht: (2026) -
Reimagining RDMA Through the Lens of ML
von: Warraich, Ertza, et al.
Veröffentlicht: (2025) -
On the Resilience of Fast Failover Routing Against Dynamic Link Failures
von: Dai, Wenkai, et al.
Veröffentlicht: (2024) -
FedRDMA: Communication-Efficient Cross-Silo Federated LLM via Chunked RDMA Transmission
von: Zhang, Zeling, et al.
Veröffentlicht: (2024) -
Closing the HPC-Cloud Convergence Gap: Multi-Tenant Slingshot RDMA for Kubernetes
von: Friese, Philipp A., et al.
Veröffentlicht: (2025)