Disaggregated Memory with SmartNIC Offloading: a Case Study on Graph Processing
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wahlgren, Jacob, Schieffer, Gabin, Gokhale, Maya, Pearce, Roger, Peng, Ivy |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Communication Offloading on SmartNIC DPUs: A Quantitative Approach
par: Wahlgren, Jacob, et autres
Publié: (2026)
par: Wahlgren, Jacob, et autres
Publié: (2026)
Multi-level Memory-Centric Profiling on ARM Processors with ARM SPE
par: Miksits, Samuel, et autres
Publié: (2024)
par: Miksits, Samuel, et autres
Publié: (2024)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
par: Wahlgren, Jacob, et autres
Publié: (2025)
par: Wahlgren, Jacob, et autres
Publié: (2025)
Inter-APU Communication on AMD MI300A Systems via Infinity Fabric: a Deep Dive
par: Schieffer, Gabin, et autres
Publié: (2025)
par: Schieffer, Gabin, et autres
Publié: (2025)
Understanding Layered Portability from HPC to Cloud in Containerized Environments
par: Medeiros, Daniel, et autres
Publié: (2024)
par: Medeiros, Daniel, et autres
Publié: (2024)
A GPU-accelerated Molecular Docking Workflow with Kubernetes and Apache Airflow
par: Medeiros, Daniel, et autres
Publié: (2024)
par: Medeiros, Daniel, et autres
Publié: (2024)
Kub: Enabling Elastic HPC Workloads on Containerized Environments
par: Medeiros, Daniel, et autres
Publié: (2024)
par: Medeiros, Daniel, et autres
Publié: (2024)
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
par: Schieffer, Gabin, et autres
Publié: (2024)
par: Schieffer, Gabin, et autres
Publié: (2024)
Accelerating Drug Discovery in AutoDock-GPU with Tensor Cores
par: Schieffer, Gabin, et autres
Publié: (2024)
par: Schieffer, Gabin, et autres
Publié: (2024)
Taming GPU Underutilization via Static Partitioning and Fine-grained CPU Offloading
par: Schieffer, Gabin, et autres
Publié: (2026)
par: Schieffer, Gabin, et autres
Publié: (2026)
ARM SVE Unleashed: Performance and Insights Across HPC Applications on Nvidia Grace
par: Shi, Ruimin, et autres
Publié: (2025)
par: Shi, Ruimin, et autres
Publié: (2025)
High-performance Vector-length Agnostic Quantum Circuit Simulations on ARM Processors
par: Shi, Ruimin, et autres
Publié: (2026)
par: Shi, Ruimin, et autres
Publié: (2026)
Plug & Offload: Transparently Offloading TCP Stack onto Off-path SmartNIC with PnO-TCP
par: Nan, Hailong, et autres
Publié: (2025)
par: Nan, Hailong, et autres
Publié: (2025)
Harnessing CUDA-Q's MPS for Tensor Network Simulations of Large-Scale Quantum Circuits
par: Schieffer, Gabin, et autres
Publié: (2025)
par: Schieffer, Gabin, et autres
Publié: (2025)
Meili: Enabling SmartNIC as a Service in the Cloud
par: Su, Qiang, et autres
Publié: (2023)
par: Su, Qiang, et autres
Publié: (2023)
The Forward-In-Time-Only Assumption in SmartNIC Resource Management: A Critique of Wave and the Case for Bilateral Interaction
par: Borrill, Paul
Publié: (2026)
par: Borrill, Paul
Publié: (2026)
OpenCUBE: Building an Open Source Cloud Blueprint with EPI Systems
par: Peng, Ivy, et autres
Publié: (2024)
par: Peng, Ivy, et autres
Publié: (2024)
Understanding Data Movement in AMD Multi-GPU Systems with Infinity Fabric
par: Schieffer, Gabin, et autres
Publié: (2024)
par: Schieffer, Gabin, et autres
Publié: (2024)
SCENIC: Stream Computation-Enhanced SmartNIC
par: Ramhorst, Benjamin, et autres
Publié: (2026)
par: Ramhorst, Benjamin, et autres
Publié: (2026)
Reliable Replication Protocols on SmartNICs
par: Katebzadeh, M. R. Siavash, et autres
Publié: (2025)
par: Katebzadeh, M. R. Siavash, et autres
Publié: (2025)
ARC-V: Vertical Resource Adaptivity for HPC Workloads in Containerized Environments
par: Medeiros, Daniel, et autres
Publié: (2025)
par: Medeiros, Daniel, et autres
Publié: (2025)
FlexKV: Flexible Index Offloading for Memory-Disaggregated Key-Value Store
par: Hu, Zhisheng, et autres
Publié: (2025)
par: Hu, Zhisheng, et autres
Publié: (2025)
Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors
par: Shi, Ruimin, et autres
Publié: (2026)
par: Shi, Ruimin, et autres
Publié: (2026)
OffloadFS: Leveraging Disaggregated Storage for Computation Offloading
par: Moon, Sungho, et autres
Publié: (2026)
par: Moon, Sungho, et autres
Publié: (2026)
DecLock: A Case of Decoupled Locking for Disaggregated Memory
par: Zhang, Hanze, et autres
Publié: (2025)
par: Zhang, Hanze, et autres
Publié: (2025)
Employ SmartNICs' Data Path Accelerators for Ordered Key-Value Stores
par: Schimmelpfennig, Frederic, et autres
Publié: (2026)
par: Schimmelpfennig, Frederic, et autres
Publié: (2026)
Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC
par: Siavashi, Mohammad, et autres
Publié: (2026)
par: Siavashi, Mohammad, et autres
Publié: (2026)
A Chronological Analysis of the Evolution of SmartNICs
par: Ajayi, Olasupo, et autres
Publié: (2025)
par: Ajayi, Olasupo, et autres
Publié: (2025)
DRackSim: Simulator for Rack-scale Memory Disaggregation
par: Puri, Amit, et autres
Publié: (2023)
par: Puri, Amit, et autres
Publié: (2023)
SWARM: Replicating Shared Disaggregated-Memory Data in No Time
par: Murat, Antoine, et autres
Publié: (2024)
par: Murat, Antoine, et autres
Publié: (2024)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
par: Ekelund, Jonah, et autres
Publié: (2025)
par: Ekelund, Jonah, et autres
Publié: (2025)
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
par: Patke, Archit, et autres
Publié: (2025)
par: Patke, Archit, et autres
Publié: (2025)
GriNNder: Breaking the Memory Capacity Wall in Full-Graph GNN Training with Storage Offloading
par: Song, Jaeyong, et autres
Publié: (2026)
par: Song, Jaeyong, et autres
Publié: (2026)
Lotus: Optimizing Disaggregated Transactions with Disaggregated Locks
par: Hu, Zhisheng, et autres
Publié: (2025)
par: Hu, Zhisheng, et autres
Publié: (2025)
PULSE: Accelerating Distributed Pointer-Traversals on Disaggregated Memory (Extended Version)
par: Tang, Yupeng, et autres
Publié: (2023)
par: Tang, Yupeng, et autres
Publié: (2023)
To Offload or Not To Offload: Model-driven Comparison of Edge-native and On-device Processing In the Era of Accelerators
par: Ng, Nathan, et autres
Publié: (2025)
par: Ng, Nathan, et autres
Publié: (2025)
Automatic BLAS Offloading on Unified Memory Architecture: A Study on NVIDIA Grace-Hopper
par: Li, Junjie, et autres
Publié: (2024)
par: Li, Junjie, et autres
Publié: (2024)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
par: Hu, Cunchen, et autres
Publié: (2024)
par: Hu, Cunchen, et autres
Publié: (2024)
DOLMA: A Data Object Level Memory Disaggregation Framework for HPC Applications
par: Zheng, Haoyu, et autres
Publié: (2025)
par: Zheng, Haoyu, et autres
Publié: (2025)
uBFT: Microsecond-scale BFT using Disaggregated Memory [Extended Version]
par: Aguilera, Marcos K., et autres
Publié: (2022)
par: Aguilera, Marcos K., et autres
Publié: (2022)
Documents similaires
-
Communication Offloading on SmartNIC DPUs: A Quantitative Approach
par: Wahlgren, Jacob, et autres
Publié: (2026) -
Multi-level Memory-Centric Profiling on ARM Processors with ARM SPE
par: Miksits, Samuel, et autres
Publié: (2024) -
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
par: Wahlgren, Jacob, et autres
Publié: (2025) -
Inter-APU Communication on AMD MI300A Systems via Infinity Fabric: a Deep Dive
par: Schieffer, Gabin, et autres
Publié: (2025) -
Understanding Layered Portability from HPC to Cloud in Containerized Environments
par: Medeiros, Daniel, et autres
Publié: (2024)