DRust: Language-Guided Distributed Shared Memory with Fine Granularity, Full Transparency, and Ultra Efficiency
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Haoran, Qiao, Yifan, Liu, Shi, Yu, Shan, Ni, Yuanjiang, Lu, Qingda, Wu, Jiesheng, Zhang, Yiying, Kim, Miryung, Xu, Harry |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
by: Qiao, Yifan, et al.
Published: (2024)
by: Qiao, Yifan, et al.
Published: (2024)
A Tale of Two Paths: Toward a Hybrid Data Plane for Efficient Far-Memory Applications
by: Chen, Lei, et al.
Published: (2024)
by: Chen, Lei, et al.
Published: (2024)
Shared Memory-Aware Latency-Sensitive Message Aggregation for Fine-Grained Communication
by: Chandrasekar, Kavitha, et al.
Published: (2024)
by: Chandrasekar, Kavitha, et al.
Published: (2024)
CXL Shared Memory Programming: Barely Distributed and Almost Persistent
by: Xu, Yi, et al.
Published: (2024)
by: Xu, Yi, et al.
Published: (2024)
Distributing Context-Aware Shared Memory Data Structures: A Case Study on Singly-Linked Lists
by: Ravishankar, Raaghav, et al.
Published: (2024)
by: Ravishankar, Raaghav, et al.
Published: (2024)
COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems
by: Raju, Aditi, et al.
Published: (2025)
by: Raju, Aditi, et al.
Published: (2025)
GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
by: Guo, Cong, et al.
Published: (2024)
by: Guo, Cong, et al.
Published: (2024)
PATSMA: Parameter Auto-tuning for Shared Memory Algorithms
by: Fernandes, Joao B., et al.
Published: (2024)
by: Fernandes, Joao B., et al.
Published: (2024)
Byzantine-Tolerant Consensus in GPU-Inspired Shared Memory
by: Georgiou, Chryssis, et al.
Published: (2025)
by: Georgiou, Chryssis, et al.
Published: (2025)
SWARM: Replicating Shared Disaggregated-Memory Data in No Time
by: Murat, Antoine, et al.
Published: (2024)
by: Murat, Antoine, et al.
Published: (2024)
Granular Synchrony
by: Giridharan, Neil, et al.
Published: (2024)
by: Giridharan, Neil, et al.
Published: (2024)
High-Performance Sorting-Based k-mer Counting in Distributed Memory with Flexible Hybrid Parallelism
by: Li, Yifan, et al.
Published: (2024)
by: Li, Yifan, et al.
Published: (2024)
MemFine: Memory-Aware Fine-Grained Scheduling for MoE Training
by: Zhao, Lu, et al.
Published: (2025)
by: Zhao, Lu, et al.
Published: (2025)
Shared Randomness Helps with Local Distributed Problems
by: Balliu, Alkida, et al.
Published: (2024)
by: Balliu, Alkida, et al.
Published: (2024)
Shared Virtual Memory: Its Design and Performance Implications for Diverse Applications
by: Cooper, Bennett, et al.
Published: (2024)
by: Cooper, Bennett, et al.
Published: (2024)
An Asynchronous Many-Task Algorithm for Unstructured $S_{N}$ Transport on Shared Memory Systems
by: Elwood, Alex, et al.
Published: (2025)
by: Elwood, Alex, et al.
Published: (2025)
EaCO: Resource Sharing Dynamics and Its Impact on Energy Efficiency for DNN Training
by: Haghshenas, Kawsar, et al.
Published: (2024)
by: Haghshenas, Kawsar, et al.
Published: (2024)
VDCores: Resource Decoupled Programming and Execution for Asynchronous GPU
by: He, Zijian, et al.
Published: (2026)
by: He, Zijian, et al.
Published: (2026)
Understanding Read-Write Wait-Free Coverings in the Fully-Anonymous Shared-Memory Model
by: Losa, Giuliano, et al.
Published: (2024)
by: Losa, Giuliano, et al.
Published: (2024)
TeraNoC: A Multi-Channel 32-bit Fine-Grained, Hybrid Mesh-Crossbar NoC for Efficient Scale-up of 1000+ Core Shared-L1-Memory Clusters
by: Zhang, Yichao, et al.
Published: (2025)
by: Zhang, Yichao, et al.
Published: (2025)
GriNNder: Breaking the Memory Capacity Wall in Full-Graph GNN Training with Storage Offloading
by: Song, Jaeyong, et al.
Published: (2026)
by: Song, Jaeyong, et al.
Published: (2026)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
by: Yoon, Dongha, et al.
Published: (2025)
by: Yoon, Dongha, et al.
Published: (2025)
Zenix: Efficient Execution of Bulky Serverless Applications
by: Guo, Zhiyuan, et al.
Published: (2022)
by: Guo, Zhiyuan, et al.
Published: (2022)
Pythia: Exploiting Workflow Predictability for Efficient Agent-Native LLM Serving
by: Yu, Shan, et al.
Published: (2026)
by: Yu, Shan, et al.
Published: (2026)
MemAscend: System Memory Optimization for SSD-Offloaded LLM Fine-Tuning
by: Liaw, Yong-Cheng, et al.
Published: (2025)
by: Liaw, Yong-Cheng, et al.
Published: (2025)
A Granularity Characterization of Task Scheduling Effectiveness
by: Anvari, Sana Taghipour, et al.
Published: (2026)
by: Anvari, Sana Taghipour, et al.
Published: (2026)
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
by: Patke, Archit, et al.
Published: (2025)
by: Patke, Archit, et al.
Published: (2025)
TrEnv-X: Transparently Share Serverless Execution Environments Across Different Functions and Nodes
by: Huang, Jialiang, et al.
Published: (2025)
by: Huang, Jialiang, et al.
Published: (2025)
Optimizing Memory Allocation in Distributed Clusters with Predictive Modeling
by: Bader, Jonathan, et al.
Published: (2026)
by: Bader, Jonathan, et al.
Published: (2026)
Solutions for Distributed Memory Access Mechanism on HPC Clusters
by: Meizner, Jan, et al.
Published: (2025)
by: Meizner, Jan, et al.
Published: (2025)
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism
by: Ai, Xin, et al.
Published: (2024)
by: Ai, Xin, et al.
Published: (2024)
DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference
by: Lin, Shouxu, et al.
Published: (2026)
by: Lin, Shouxu, et al.
Published: (2026)
VersaSlot: Efficient Fine-grained FPGA Sharing with Big.Little Slots and Live Migration in FPGA Cluster
by: Gu, Jianfeng, et al.
Published: (2025)
by: Gu, Jianfeng, et al.
Published: (2025)
Analysis and Optimized CXL-Attached Memory Allocation for Long-Context LLM Fine-Tuning
by: Liaw, Yong-Cheng, et al.
Published: (2025)
by: Liaw, Yong-Cheng, et al.
Published: (2025)
Memory-Efficient Split Federated Learning for LLM Fine-Tuning on Heterogeneous Mobile Devices
by: Chen, Xiaopei, et al.
Published: (2025)
by: Chen, Xiaopei, et al.
Published: (2025)
Memory-Efficient Federated Fine-Tuning of Large Language Models via Layer Pruning
by: Wu, Yebo, et al.
Published: (2025)
by: Wu, Yebo, et al.
Published: (2025)
Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
by: Srivatsa, Vikranth, et al.
Published: (2026)
by: Srivatsa, Vikranth, et al.
Published: (2026)
Self-Evolving Distributed Memory Architecture for Scalable AI Systems
by: Li, Zixuan, et al.
Published: (2026)
by: Li, Zixuan, et al.
Published: (2026)
On the Performance and Memory Footprint of Distributed Training: An Empirical Study on Transformers
by: Lu, Zhengxian, et al.
Published: (2024)
by: Lu, Zhengxian, et al.
Published: (2024)
Efficient and Portable Support for Overdecomposition on Distributed Memory GPGPU Platforms
by: Bhosale, Aditya, et al.
Published: (2026)
by: Bhosale, Aditya, et al.
Published: (2026)
Similar Items
-
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
by: Qiao, Yifan, et al.
Published: (2024) -
A Tale of Two Paths: Toward a Hybrid Data Plane for Efficient Far-Memory Applications
by: Chen, Lei, et al.
Published: (2024) -
Shared Memory-Aware Latency-Sensitive Message Aggregation for Fine-Grained Communication
by: Chandrasekar, Kavitha, et al.
Published: (2024) -
CXL Shared Memory Programming: Barely Distributed and Almost Persistent
by: Xu, Yi, et al.
Published: (2024) -
Distributing Context-Aware Shared Memory Data Structures: A Case Study on Singly-Linked Lists
by: Ravishankar, Raaghav, et al.
Published: (2024)