An Extensible Software Transport Layer for GPU Networking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Yang, Chen, Zhongjie, Mao, Ziming, Lao, ChonLam, Yang, Shuo, Kannan, Pravein Govindan, Gao, Jiaqi, Zhao, Yilong, Wu, Yongji, You, Kaichao, Ren, Fengyuan, Xu, Zhiying, Raiciu, Costin, Stoica, Ion |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UCCL-EP: Portable Expert-Parallel Communication
von: Mao, Ziming, et al.
Veröffentlicht: (2025)
von: Mao, Ziming, et al.
Veröffentlicht: (2025)
UCCL-Zip: Lossless Compression Supercharged GPU Communication
von: Ma, Shuang, et al.
Veröffentlicht: (2026)
von: Ma, Shuang, et al.
Veröffentlicht: (2026)
EdgeSight: Enabling Modeless and Cost-Efficient Inference at the Edge
von: Lao, ChonLam, et al.
Veröffentlicht: (2024)
von: Lao, ChonLam, et al.
Veröffentlicht: (2024)
Towards Easy and Realistic Network Infrastructure Testing for Large-scale Machine Learning
von: Yoo, Jinsun, et al.
Veröffentlicht: (2025)
von: Yoo, Jinsun, et al.
Veröffentlicht: (2025)
THC: Accelerating Distributed Deep Learning Using Tensor Homomorphic Compression
von: Li, Minghao, et al.
Veröffentlicht: (2023)
von: Li, Minghao, et al.
Veröffentlicht: (2023)
LEANN: A Low-Storage Vector Index
von: Wang, Yichuan, et al.
Veröffentlicht: (2025)
von: Wang, Yichuan, et al.
Veröffentlicht: (2025)
Prose-to-P4: Leveraging High Level Languages
von: Dumitru, Mihai-Valentin, et al.
Veröffentlicht: (2024)
von: Dumitru, Mihai-Valentin, et al.
Veröffentlicht: (2024)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
von: Qiao, Yifan, et al.
Veröffentlicht: (2024)
von: Qiao, Yifan, et al.
Veröffentlicht: (2024)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
von: Jiang, Xuanlin, et al.
Veröffentlicht: (2024)
von: Jiang, Xuanlin, et al.
Veröffentlicht: (2024)
Some Present-Day Problems of Romanian Library Science
von: Stoica, Ion
Veröffentlicht: (1973)
von: Stoica, Ion
Veröffentlicht: (1973)
The Central University Library, Bucharest. Over Seventy-five Years in the History of a Collection
von: Stoica, Ion
Veröffentlicht: (1972)
von: Stoica, Ion
Veröffentlicht: (1972)
depyf: Open the Opaque Box of PyTorch Compiler for Machine Learning Researchers
von: You, Kaichao, et al.
Veröffentlicht: (2024)
von: You, Kaichao, et al.
Veröffentlicht: (2024)
Revisiting Cache Freshness for Emerging Real-Time Applications
von: Mao, Ziming, et al.
Veröffentlicht: (2024)
von: Mao, Ziming, et al.
Veröffentlicht: (2024)
BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching
von: Zhao, Yilong, et al.
Veröffentlicht: (2024)
von: Zhao, Yilong, et al.
Veröffentlicht: (2024)
K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model
von: Cao, Shiyi, et al.
Veröffentlicht: (2026)
von: Cao, Shiyi, et al.
Veröffentlicht: (2026)
Pie: Pooling CPU Memory for LLM Inference
von: Xu, Yi, et al.
Veröffentlicht: (2024)
von: Xu, Yi, et al.
Veröffentlicht: (2024)
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start
von: Liu, Xueshen, et al.
Veröffentlicht: (2026)
von: Liu, Xueshen, et al.
Veröffentlicht: (2026)
Boosting green innovation on corporate performance: Managerial environmental concern's moderating role
von: Thanh Tiep Le, et al.
Veröffentlicht: (2024)
von: Thanh Tiep Le, et al.
Veröffentlicht: (2024)
Post-Training Sparse Attention with Double Sparsity
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
Towards Efficient and Practical GPU Multitasking in the Era of LLM
von: Xing, Jiarong, et al.
Veröffentlicht: (2025)
von: Xing, Jiarong, et al.
Veröffentlicht: (2025)
LLMic: Romanian Foundation Language Model
von: Bădoiu, Vlad-Andrei, et al.
Veröffentlicht: (2025)
von: Bădoiu, Vlad-Andrei, et al.
Veröffentlicht: (2025)
FuLG: 150B Romanian Corpus for Language Model Pretraining
von: Bădoiu, Vlad-Andrei, et al.
Veröffentlicht: (2024)
von: Bădoiu, Vlad-Andrei, et al.
Veröffentlicht: (2024)
A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM
von: Xi, Shaoke, et al.
Veröffentlicht: (2026)
von: Xi, Shaoke, et al.
Veröffentlicht: (2026)
Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
von: Yu, Shan, et al.
Veröffentlicht: (2025)
von: Yu, Shan, et al.
Veröffentlicht: (2025)
Unveiling the value of institutional pressure in socially sustainable supply chain management: The role of top management support for social initiatives and organisational culture
von: Min Zhou, et al.
Veröffentlicht: (2024)
von: Min Zhou, et al.
Veröffentlicht: (2024)
HashAttention: Semantic Sparsity for Faster Inference
von: Desai, Aditya, et al.
Veröffentlicht: (2024)
von: Desai, Aditya, et al.
Veröffentlicht: (2024)
GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents
von: Shetty, Manish, et al.
Veröffentlicht: (2025)
von: Shetty, Manish, et al.
Veröffentlicht: (2025)
Mélange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
von: Griggs, Tyler, et al.
Veröffentlicht: (2024)
von: Griggs, Tyler, et al.
Veröffentlicht: (2024)
TrainMover: An Interruption-Resilient Runtime for ML Training
von: Lao, ChonLam, et al.
Veröffentlicht: (2024)
von: Lao, ChonLam, et al.
Veröffentlicht: (2024)
RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
GPU Performance Portability needs Autotuning
von: Ringlein, Burkhard, et al.
Veröffentlicht: (2025)
von: Ringlein, Burkhard, et al.
Veröffentlicht: (2025)
Unleashing Scalable Context Parallelism for Foundation Models Pre-Training via FCP
von: Zhao, Yilong, et al.
Veröffentlicht: (2026)
von: Zhao, Yilong, et al.
Veröffentlicht: (2026)
Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
I've Got 99 Problems But FLOPS Ain't One
von: Gherghescu, Alexandru M., et al.
Veröffentlicht: (2024)
von: Gherghescu, Alexandru M., et al.
Veröffentlicht: (2024)
Transportation of fish in India: problems and prospects
von: Peregreen, P.A., et al.
Veröffentlicht: (1969)
von: Peregreen, P.A., et al.
Veröffentlicht: (1969)
FAIRSECO: An Extensible Framework for Impact Measurement of Research Software
von: Deekshitha, et al.
Veröffentlicht: (2024)
von: Deekshitha, et al.
Veröffentlicht: (2024)
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
SkyServe: Serving AI Models across Regions and Clouds with Spot Instances
von: Mao, Ziming, et al.
Veröffentlicht: (2024)
von: Mao, Ziming, et al.
Veröffentlicht: (2024)
SMaRTT: Sender-based Marked Rapidly-adapting Trimmed & Timed Transport
von: Bonato, Tommaso, et al.
Veröffentlicht: (2024)
von: Bonato, Tommaso, et al.
Veröffentlicht: (2024)
Towards Ideal Temporal Graph Neural Networks: Evaluations and Conclusions after 10,000 GPU Hours
von: Yang, Yuxin, et al.
Veröffentlicht: (2024)
von: Yang, Yuxin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
UCCL-EP: Portable Expert-Parallel Communication
von: Mao, Ziming, et al.
Veröffentlicht: (2025) -
UCCL-Zip: Lossless Compression Supercharged GPU Communication
von: Ma, Shuang, et al.
Veröffentlicht: (2026) -
EdgeSight: Enabling Modeless and Cost-Efficient Inference at the Edge
von: Lao, ChonLam, et al.
Veröffentlicht: (2024) -
Towards Easy and Realistic Network Infrastructure Testing for Large-scale Machine Learning
von: Yoo, Jinsun, et al.
Veröffentlicht: (2025) -
THC: Accelerating Distributed Deep Learning Using Tensor Homomorphic Compression
von: Li, Minghao, et al.
Veröffentlicht: (2023)