fabric-lib: RDMA Point-to-Point Communication for LLM Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Licker, Nandor, Hu, Kevin, Zaytsev, Vladimir, Chen, Lequn |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ALock: Asymmetric Lock Primitive for RDMA Systems
di: Baran, Amanda, et al.
Pubblicazione: (2024)
di: Baran, Amanda, et al.
Pubblicazione: (2024)
Towards Efficient and Scalable Distributed Vector Search with RDMA
di: Zhi, Xiangyu, et al.
Pubblicazione: (2025)
di: Zhi, Xiangyu, et al.
Pubblicazione: (2025)
Portable, heterogeneous ensemble workflows at scale using libEnsemble
di: Hudson, Stephen, et al.
Pubblicazione: (2024)
di: Hudson, Stephen, et al.
Pubblicazione: (2024)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
di: Brock, Benjamin, et al.
Pubblicazione: (2023)
di: Brock, Benjamin, et al.
Pubblicazione: (2023)
The Semantic Arrow of Time, Part III: RDMA and the Completion Fallacy
di: Borrill, Paul
Pubblicazione: (2026)
di: Borrill, Paul
Pubblicazione: (2026)
OnePiece: A Large-Scale Distributed Inference System with RDMA for Complex AI-Generated Content (AIGC) Workflows
di: Chen, June, et al.
Pubblicazione: (2026)
di: Chen, June, et al.
Pubblicazione: (2026)
FedRDMA: Communication-Efficient Cross-Silo Federated LLM via Chunked RDMA Transmission
di: Zhang, Zeling, et al.
Pubblicazione: (2024)
di: Zhang, Zeling, et al.
Pubblicazione: (2024)
On-the-fly Communication-and-Computing to Enable Representation Learning for Distributed Point Clouds
di: Chen, Xu, et al.
Pubblicazione: (2024)
di: Chen, Xu, et al.
Pubblicazione: (2024)
Reimagining RDMA Through the Lens of ML
di: Warraich, Ertza, et al.
Pubblicazione: (2025)
di: Warraich, Ertza, et al.
Pubblicazione: (2025)
Handling of Memory Page Faults during Virtual-Address RDMA
di: Psistakis, Antonis
Pubblicazione: (2025)
di: Psistakis, Antonis
Pubblicazione: (2025)
On the Universality of Round Elimination Fixed Points
di: Balliu, Alkida, et al.
Pubblicazione: (2025)
di: Balliu, Alkida, et al.
Pubblicazione: (2025)
AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training
di: Chen, Qiaoling, et al.
Pubblicazione: (2023)
di: Chen, Qiaoling, et al.
Pubblicazione: (2023)
Thallus: An RDMA-based Columnar Data Transport Protocol
di: Chakraborty, Jayjeet, et al.
Pubblicazione: (2024)
di: Chakraborty, Jayjeet, et al.
Pubblicazione: (2024)
Varuna: Enabling Failure-Type Aware RDMA Failover
di: Wang, Xiaoyang, et al.
Pubblicazione: (2026)
di: Wang, Xiaoyang, et al.
Pubblicazione: (2026)
Federated Inference for Heterogeneous LLM Communication and Collaboration
di: Chen, Zihan, et al.
Pubblicazione: (2026)
di: Chen, Zihan, et al.
Pubblicazione: (2026)
NestedFP: High-Performance, Memory-Efficient Dual-Precision Floating Point Support for LLMs
di: Lee, Haeun, et al.
Pubblicazione: (2025)
di: Lee, Haeun, et al.
Pubblicazione: (2025)
Fast Topology-Aware Lossy Data Compression with Full Preservation of Critical Points and Local Order
di: Fallin, Alex, et al.
Pubblicazione: (2026)
di: Fallin, Alex, et al.
Pubblicazione: (2026)
SDSL-Solver: Scalable Distributed Sparse Linear Solvers for Large-Scale Interior Point Methods
di: Yang, Shaofeng, et al.
Pubblicazione: (2026)
di: Yang, Shaofeng, et al.
Pubblicazione: (2026)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
di: Chen, Jiu, et al.
Pubblicazione: (2026)
di: Chen, Jiu, et al.
Pubblicazione: (2026)
gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
Floating-Point Data Transformation for Lossless Compression
di: Jamalidinan, Samirasadat, et al.
Pubblicazione: (2025)
di: Jamalidinan, Samirasadat, et al.
Pubblicazione: (2025)
Specifying and Verifying RDMA Synchronisation (Extended Version)
di: Ambal, Guillaume, et al.
Pubblicazione: (2026)
di: Ambal, Guillaume, et al.
Pubblicazione: (2026)
Closing the HPC-Cloud Convergence Gap: Multi-Tenant Slingshot RDMA for Kubernetes
di: Friese, Philipp A., et al.
Pubblicazione: (2025)
di: Friese, Philipp A., et al.
Pubblicazione: (2025)
Towards Federated Learning with On-device Training and Communication in 8-bit Floating Point
di: Wang, Bokun, et al.
Pubblicazione: (2024)
di: Wang, Bokun, et al.
Pubblicazione: (2024)
MemAscend: System Memory Optimization for SSD-Offloaded LLM Fine-Tuning
di: Liaw, Yong-Cheng, et al.
Pubblicazione: (2025)
di: Liaw, Yong-Cheng, et al.
Pubblicazione: (2025)
Communication-Efficient Collaborative LLM Inference over LEO Satellite Networks
di: Zhang, Songge, et al.
Pubblicazione: (2026)
di: Zhang, Songge, et al.
Pubblicazione: (2026)
Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
di: Xu, Guanbin, et al.
Pubblicazione: (2026)
di: Xu, Guanbin, et al.
Pubblicazione: (2026)
PVU: Design and Implementation of a Posit Vector Arithmetic Unit (PVU) for Enhanced Floating-Point Computing in Edge and AI Applications
di: Wu, Xinyu, et al.
Pubblicazione: (2025)
di: Wu, Xinyu, et al.
Pubblicazione: (2025)
OptiNIC: A Resilient and Tail-Optimal RDMA NIC for Distributed ML Workloads
di: Warraich, Ertza, et al.
Pubblicazione: (2025)
di: Warraich, Ertza, et al.
Pubblicazione: (2025)
Cloud Native System for LLM Inference Serving
di: Xu, Minxian, et al.
Pubblicazione: (2025)
di: Xu, Minxian, et al.
Pubblicazione: (2025)
Symphony: Optimized DNN Model Serving using Deferred Batch Scheduling
di: Chen, Lequn, et al.
Pubblicazione: (2023)
di: Chen, Lequn, et al.
Pubblicazione: (2023)
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
di: Bai, Fengyao, et al.
Pubblicazione: (2026)
di: Bai, Fengyao, et al.
Pubblicazione: (2026)
Hiding Communication Cost in Distributed LLM Training via Micro-batch Co-execution
di: Wang, Haiquan, et al.
Pubblicazione: (2024)
di: Wang, Haiquan, et al.
Pubblicazione: (2024)
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
di: Wilkins, Grant, et al.
Pubblicazione: (2024)
di: Wilkins, Grant, et al.
Pubblicazione: (2024)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
Three Birds, One Stone: Solving the Communication-Memory-Privacy Trilemma in LLM Fine-tuning Over Wireless Networks with Zeroth-Order Optimization
di: Cai, Zhijie, et al.
Pubblicazione: (2026)
di: Cai, Zhijie, et al.
Pubblicazione: (2026)
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production
di: Xue, Chunyu, et al.
Pubblicazione: (2026)
di: Xue, Chunyu, et al.
Pubblicazione: (2026)
OmniInfer: System-Wide Acceleration Techniques for Optimizing LLM Serving Throughput and Latency
di: Wang, Jun, et al.
Pubblicazione: (2025)
di: Wang, Jun, et al.
Pubblicazione: (2025)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
CrashEventLLM: Predicting System Crashes with Large Language Models
di: Mudgal, Priyanka, et al.
Pubblicazione: (2024)
di: Mudgal, Priyanka, et al.
Pubblicazione: (2024)
Documenti analoghi
-
ALock: Asymmetric Lock Primitive for RDMA Systems
di: Baran, Amanda, et al.
Pubblicazione: (2024) -
Towards Efficient and Scalable Distributed Vector Search with RDMA
di: Zhi, Xiangyu, et al.
Pubblicazione: (2025) -
Portable, heterogeneous ensemble workflows at scale using libEnsemble
di: Hudson, Stephen, et al.
Pubblicazione: (2024) -
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
di: Brock, Benjamin, et al.
Pubblicazione: (2023) -
The Semantic Arrow of Time, Part III: RDMA and the Completion Fallacy
di: Borrill, Paul
Pubblicazione: (2026)