Reimagining RDMA Through the Lens of ML
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Warraich, Ertza, Imran, Ali, Zulfiqar, Annus, Vargaftik, Shay, Fahmy, Sonia, Shahbaz, Muhammad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OptiNIC: A Resilient and Tail-Optimal RDMA NIC for Distributed ML Workloads
von: Warraich, Ertza, et al.
Veröffentlicht: (2025)
von: Warraich, Ertza, et al.
Veröffentlicht: (2025)
OptiReduce: Resilient and Tail-Optimal AllReduce for Distributed Deep Learning in the Cloud
von: Warraich, Ertza, et al.
Veröffentlicht: (2023)
von: Warraich, Ertza, et al.
Veröffentlicht: (2023)
Varuna: Enabling Failure-Type Aware RDMA Failover
von: Wang, Xiaoyang, et al.
Veröffentlicht: (2026)
von: Wang, Xiaoyang, et al.
Veröffentlicht: (2026)
Closing the HPC-Cloud Convergence Gap: Multi-Tenant Slingshot RDMA for Kubernetes
von: Friese, Philipp A., et al.
Veröffentlicht: (2025)
von: Friese, Philipp A., et al.
Veröffentlicht: (2025)
Palladium: A DPU-enabled Multi-Tenant Serverless Cloud over Zero-copy Multi-node RDMA Fabrics
von: Qi, Shixiong, et al.
Veröffentlicht: (2025)
von: Qi, Shixiong, et al.
Veröffentlicht: (2025)
FedRDMA: Communication-Efficient Cross-Silo Federated LLM via Chunked RDMA Transmission
von: Zhang, Zeling, et al.
Veröffentlicht: (2024)
von: Zhang, Zeling, et al.
Veröffentlicht: (2024)
Enabling Blockchain Interoperability Through Network Discovery Services
von: Hassan, Khalid, et al.
Veröffentlicht: (2025)
von: Hassan, Khalid, et al.
Veröffentlicht: (2025)
Runtime Verification Containers for Publish/Subscribe Networks
von: Mehran, Ali, et al.
Veröffentlicht: (2024)
von: Mehran, Ali, et al.
Veröffentlicht: (2024)
LIMO: Load-balanced Offloading with MAPE and Particle Swarm Optimization in Mobile Fog Networks
von: Seraj, Yasaman, et al.
Veröffentlicht: (2024)
von: Seraj, Yasaman, et al.
Veröffentlicht: (2024)
Lightweight Latency Prediction Scheme for Edge Applications: A Rational Modelling Approach
von: Liyanage, Mohan, et al.
Veröffentlicht: (2025)
von: Liyanage, Mohan, et al.
Veröffentlicht: (2025)
EdgeLinker: Practical Blockchain-based Framework for Healthcare Fog Applications to Enhance Security in Edge-IoT Data Communications
von: Zarkesh, Mahdi Akbari, et al.
Veröffentlicht: (2024)
von: Zarkesh, Mahdi Akbari, et al.
Veröffentlicht: (2024)
Graph neural network for in-network placement of real-time metaverse tasks in next-generation network
von: Rashid, Sulaiman Muhammad, et al.
Veröffentlicht: (2024)
von: Rashid, Sulaiman Muhammad, et al.
Veröffentlicht: (2024)
Joint wireless and computing resource management with optimal slice selection in in-network-edge metaverse system
von: Rashid, Sulaiman Muhammad, et al.
Veröffentlicht: (2024)
von: Rashid, Sulaiman Muhammad, et al.
Veröffentlicht: (2024)
FlowTracer: A Tool for Uncovering Network Path Usage Imbalance in AI Training Clusters
von: Jamil, Hasibul, et al.
Veröffentlicht: (2024)
von: Jamil, Hasibul, et al.
Veröffentlicht: (2024)
Optimizing Split Learning Latency in TinyML-Based IoT Systems
von: Jenhani, Zied, et al.
Veröffentlicht: (2025)
von: Jenhani, Zied, et al.
Veröffentlicht: (2025)
DynamiQ: Accelerating Gradient Synchronization using Compressed Multi-hop All-reduce
von: Han, Wenchen, et al.
Veröffentlicht: (2026)
von: Han, Wenchen, et al.
Veröffentlicht: (2026)
A Uniqueness Theorem for Distributed Computation under Physical Constraint
von: Ren, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Ren, Zhiyuan, et al.
Veröffentlicht: (2025)
Fast Multichannel Topology Discovery in Cognitive Radio Networks
von: Wang, Yung-Li, et al.
Veröffentlicht: (2025)
von: Wang, Yung-Li, et al.
Veröffentlicht: (2025)
Enabling Scalability in Asynchronous and Bidirectional Communication in LPWAN
von: Rahman, Mahbubur
Veröffentlicht: (2025)
von: Rahman, Mahbubur
Veröffentlicht: (2025)
Real-Time In-Network Machine Learning on P4-Programmable FPGA SmartNICs with Fixed-Point Arithmetic and Taylor
von: Sada, Mohammad Firas, et al.
Veröffentlicht: (2025)
von: Sada, Mohammad Firas, et al.
Veröffentlicht: (2025)
Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration
von: Luo, Haoxiang, et al.
Veröffentlicht: (2025)
von: Luo, Haoxiang, et al.
Veröffentlicht: (2025)
Performance Evaluation of Brokerless Messaging Libraries
von: La Corte, Lorenzo, et al.
Veröffentlicht: (2025)
von: La Corte, Lorenzo, et al.
Veröffentlicht: (2025)
Accelerating Stable Matching between Workers and Spatial-Temporal Tasks for Dynamic MCS: A Stagewise Service Trading Approach
von: Qi, Houyi, et al.
Veröffentlicht: (2025)
von: Qi, Houyi, et al.
Veröffentlicht: (2025)
Carbon-Aware Temporal Data Transfer Scheduling Across Cloud Datacenters
von: Rodrigues, Elvis, et al.
Veröffentlicht: (2025)
von: Rodrigues, Elvis, et al.
Veröffentlicht: (2025)
RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
von: Xu, Heng, et al.
Veröffentlicht: (2025)
von: Xu, Heng, et al.
Veröffentlicht: (2025)
PeerSync: Accelerating Containerized Service Delivery at the Network Edge
von: Deng, Yinuo, et al.
Veröffentlicht: (2025)
von: Deng, Yinuo, et al.
Veröffentlicht: (2025)
Edge-assisted Parallel Uncertain Skyline Processing for Low-latency IoE Analysis
von: Lai, Chuan-Chi, et al.
Veröffentlicht: (2025)
von: Lai, Chuan-Chi, et al.
Veröffentlicht: (2025)
QoS-Aware Load Balancing in the Computing Continuum via Multi-Player Bandits
von: Čilić, Ivan, et al.
Veröffentlicht: (2025)
von: Čilić, Ivan, et al.
Veröffentlicht: (2025)
AllReduce Scheduling with Hierarchical Deep Reinforcement Learning
von: Wei, Yufan, et al.
Veröffentlicht: (2025)
von: Wei, Yufan, et al.
Veröffentlicht: (2025)
Leveraging InfiniBand Controller to Configure Deadlock-Free Routing Engines for Dragonflies
von: Maglione-Mathey, German, et al.
Veröffentlicht: (2025)
von: Maglione-Mathey, German, et al.
Veröffentlicht: (2025)
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
von: Du, Chengze, et al.
Veröffentlicht: (2025)
von: Du, Chengze, et al.
Veröffentlicht: (2025)
Revolutionizing Datacenter Networks via Reconfigurable Topologies
von: Avin, Chen, et al.
Veröffentlicht: (2025)
von: Avin, Chen, et al.
Veröffentlicht: (2025)
MDTP -- An Adaptive Multi-Source Data Transfer Protocol
von: Abdollah, Sepideh, et al.
Veröffentlicht: (2025)
von: Abdollah, Sepideh, et al.
Veröffentlicht: (2025)
Future Resource Bank for ISAC: Achieving Fast and Stable Win-Win Matching for Both Individuals and Coalitions
von: Qi, Houyi, et al.
Veröffentlicht: (2025)
von: Qi, Houyi, et al.
Veröffentlicht: (2025)
Short-circuiting Rings for Low-Latency AllReduce
von: Hammer, Sarah-Michelle, et al.
Veröffentlicht: (2025)
von: Hammer, Sarah-Michelle, et al.
Veröffentlicht: (2025)
Toward Co-adapting Machine Learning Job Shape and Cluster Topology
von: Chen, Shawn Shuoshuo, et al.
Veröffentlicht: (2025)
von: Chen, Shawn Shuoshuo, et al.
Veröffentlicht: (2025)
POSMAC: Powering Up In-Network AR/CG Traffic Classification with Online Learning
von: Shirmarz, Alireza, et al.
Veröffentlicht: (2025)
von: Shirmarz, Alireza, et al.
Veröffentlicht: (2025)
Resource Allocation Driven by Large Models in Future Semantic-Aware Networks
von: Zhang, Haijun, et al.
Veröffentlicht: (2025)
von: Zhang, Haijun, et al.
Veröffentlicht: (2025)
FAFO: Over 1 million TPS on a single node running EVM while still Merkleizing every block
von: Zarick, Ryan, et al.
Veröffentlicht: (2025)
von: Zarick, Ryan, et al.
Veröffentlicht: (2025)
A Study on 5G Network Slice Isolation Based on Native Cloud and Edge Computing Tools
von: Andrade, Maiko, et al.
Veröffentlicht: (2025)
von: Andrade, Maiko, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
OptiNIC: A Resilient and Tail-Optimal RDMA NIC for Distributed ML Workloads
von: Warraich, Ertza, et al.
Veröffentlicht: (2025) -
OptiReduce: Resilient and Tail-Optimal AllReduce for Distributed Deep Learning in the Cloud
von: Warraich, Ertza, et al.
Veröffentlicht: (2023) -
Varuna: Enabling Failure-Type Aware RDMA Failover
von: Wang, Xiaoyang, et al.
Veröffentlicht: (2026) -
Closing the HPC-Cloud Convergence Gap: Multi-Tenant Slingshot RDMA for Kubernetes
von: Friese, Philipp A., et al.
Veröffentlicht: (2025) -
Palladium: A DPU-enabled Multi-Tenant Serverless Cloud over Zero-copy Multi-node RDMA Fabrics
von: Qi, Shixiong, et al.
Veröffentlicht: (2025)