emucxl: an emulation framework for CXL-based disaggregated memory applications
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gond, Raja, Kulkarni, Purushottam |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
von: Gond, Raja, et al.
Veröffentlicht: (2025)
von: Gond, Raja, et al.
Veröffentlicht: (2025)
Towards CXL Resilience to CPU Failures
von: Psistakis, Antonis, et al.
Veröffentlicht: (2026)
von: Psistakis, Antonis, et al.
Veröffentlicht: (2026)
OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
von: Wu, Siyu, et al.
Veröffentlicht: (2025)
von: Wu, Siyu, et al.
Veröffentlicht: (2025)
CXL Shared Memory Programming: Barely Distributed and Almost Persistent
von: Xu, Yi, et al.
Veröffentlicht: (2024)
von: Xu, Yi, et al.
Veröffentlicht: (2024)
Modeling the Potential of Message-Free Communication via CXL.mem
von: Vanecek, Stepan, et al.
Veröffentlicht: (2025)
von: Vanecek, Stepan, et al.
Veröffentlicht: (2025)
MPI-over-CXL: Enhancing Communication Efficiency in Distributed HPC Systems
von: Kwon, Miryeong, et al.
Veröffentlicht: (2025)
von: Kwon, Miryeong, et al.
Veröffentlicht: (2025)
MQFQ-Sticky: Fair Queueing For Serverless GPU Functions
von: Fuerst, Alexander, et al.
Veröffentlicht: (2025)
von: Fuerst, Alexander, et al.
Veröffentlicht: (2025)
Analysis and Optimized CXL-Attached Memory Allocation for Long-Context LLM Fine-Tuning
von: Liaw, Yong-Cheng, et al.
Veröffentlicht: (2025)
von: Liaw, Yong-Cheng, et al.
Veröffentlicht: (2025)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
von: Yoon, Dongha, et al.
Veröffentlicht: (2025)
von: Yoon, Dongha, et al.
Veröffentlicht: (2025)
ScalePool: Hybrid XLink-CXL Fabric for Composable Resource Disaggregation in Unified Scale-up Domains
von: Woo, Hyein, et al.
Veröffentlicht: (2025)
von: Woo, Hyein, et al.
Veröffentlicht: (2025)
LLM-42: Enabling Determinism in LLM Inference with Verified Speculation
von: Gond, Raja, et al.
Veröffentlicht: (2026)
von: Gond, Raja, et al.
Veröffentlicht: (2026)
The workflow motif: a widely-useful performance diagnosis abstraction for distributed applications
von: Abdi, Mania, et al.
Veröffentlicht: (2025)
von: Abdi, Mania, et al.
Veröffentlicht: (2025)
An experimental evaluation of satellite constellation emulators
von: Cionca, Victor, et al.
Veröffentlicht: (2026)
von: Cionca, Victor, et al.
Veröffentlicht: (2026)
A Programming Model for Disaggregated Memory over CXL
von: Assa, Gal, et al.
Veröffentlicht: (2024)
von: Assa, Gal, et al.
Veröffentlicht: (2024)
DPC: A Distributed Page Cache over CXL
von: Bergman, Shai, et al.
Veröffentlicht: (2026)
von: Bergman, Shai, et al.
Veröffentlicht: (2026)
FIRED: a fine-grained robust performance diagnosis framework for cloud applications
von: Xin, Ruyue, et al.
Veröffentlicht: (2022)
von: Xin, Ruyue, et al.
Veröffentlicht: (2022)
Telepathic Datacenters: Fast RPCs using Shared CXL Memory
von: Mahar, Suyash, et al.
Veröffentlicht: (2024)
von: Mahar, Suyash, et al.
Veröffentlicht: (2024)
Equilibria: Fair Multi-Tenant CXL Memory Tiering At Scale
von: Zhao, Kaiyang, et al.
Veröffentlicht: (2026)
von: Zhao, Kaiyang, et al.
Veröffentlicht: (2026)
CCCL: Node-Spanning GPU Collectives with CXL Memory Pooling
von: Xu, Dong, et al.
Veröffentlicht: (2026)
von: Xu, Dong, et al.
Veröffentlicht: (2026)
HybridTier: an Adaptive and Lightweight CXL-Memory Tiering System
von: Song, Kevin, et al.
Veröffentlicht: (2023)
von: Song, Kevin, et al.
Veröffentlicht: (2023)
Pooling Engram Conditional Memory in Large Language Models using CXL
von: Ma, Ruiyang, et al.
Veröffentlicht: (2026)
von: Ma, Ruiyang, et al.
Veröffentlicht: (2026)
Tolerance to Asynchrony of an Algorithm for Gathering Myopic Robots on an Infinite Triangular Grid
von: Gupta, Arya Tanmay, et al.
Veröffentlicht: (2023)
von: Gupta, Arya Tanmay, et al.
Veröffentlicht: (2023)
Fully Lattice-Linear Algorithms
von: Gupta, Arya Tanmay, et al.
Veröffentlicht: (2022)
von: Gupta, Arya Tanmay, et al.
Veröffentlicht: (2022)
Tolerance to Asynchrony in Algorithms for Multiplication and Modulo
von: Gupta, Arya Tanmay, et al.
Veröffentlicht: (2023)
von: Gupta, Arya Tanmay, et al.
Veröffentlicht: (2023)
Scaling atomic ordering in shared memory
von: Martignetti, Lorenzo, et al.
Veröffentlicht: (2025)
von: Martignetti, Lorenzo, et al.
Veröffentlicht: (2025)
Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management
von: Yang, Xinjun, et al.
Veröffentlicht: (2025)
von: Yang, Xinjun, et al.
Veröffentlicht: (2025)
Towards a Scalable In Situ Fast Fourier Transform
von: Kulkarni, Sudhanshu, et al.
Veröffentlicht: (2024)
von: Kulkarni, Sudhanshu, et al.
Veröffentlicht: (2024)
Asynchronous Checkpoint for Eventually Consistent Databases
von: Ravishankar, Raaghav, et al.
Veröffentlicht: (2025)
von: Ravishankar, Raaghav, et al.
Veröffentlicht: (2025)
Large Scale Multi-GPU Based Parallel Traffic Simulation for Accelerated Traffic Assignment and Propagation
von: Jiang, Xuan, et al.
Veröffentlicht: (2024)
von: Jiang, Xuan, et al.
Veröffentlicht: (2024)
A GPU accelerated mixed-precision Smoothed Particle Hydrodynamics framework with cell-based relative coordinates
von: Mao, Zirui, et al.
Veröffentlicht: (2023)
von: Mao, Zirui, et al.
Veröffentlicht: (2023)
Distributing Context-Aware Shared Memory Data Structures: A Case Study on Singly-Linked Lists
von: Ravishankar, Raaghav, et al.
Veröffentlicht: (2024)
von: Ravishankar, Raaghav, et al.
Veröffentlicht: (2024)
Transforming Lock-free Linked Lists into Distributed Lock-free Linked Lists
von: Ravishankar, Raaghav, et al.
Veröffentlicht: (2025)
von: Ravishankar, Raaghav, et al.
Veröffentlicht: (2025)
Characterizing Production GPU Workloads using System-wide Telemetry Data
von: Cankur, Onur, et al.
Veröffentlicht: (2025)
von: Cankur, Onur, et al.
Veröffentlicht: (2025)
Perpetual Exploration of a Ring in Presence of Byzantine Black Hole
von: Goswami, Pritam, et al.
Veröffentlicht: (2024)
von: Goswami, Pritam, et al.
Veröffentlicht: (2024)
Zero-consistency root emulation for unprivileged container image build
von: Priedhorsky, Reid, et al.
Veröffentlicht: (2024)
von: Priedhorsky, Reid, et al.
Veröffentlicht: (2024)
A sparsity-aware distributed-memory algorithm for sparse-sparse matrix multiplication
von: Hong, Yuxi, et al.
Veröffentlicht: (2024)
von: Hong, Yuxi, et al.
Veröffentlicht: (2024)
DUMBO: Making durable read-only transactions fly on hardware transactional memory
von: Barreto, João, et al.
Veröffentlicht: (2024)
von: Barreto, João, et al.
Veröffentlicht: (2024)
Trace-based, time-resolved analysis of MPI application performance using standard metrics
von: Haldar, Kingshuk
Veröffentlicht: (2025)
von: Haldar, Kingshuk
Veröffentlicht: (2025)
Exploring and Evaluating Real-world CXL: Use Cases and System Adoption
von: Wang, Xi, et al.
Veröffentlicht: (2024)
von: Wang, Xi, et al.
Veröffentlicht: (2024)
Scalable mRMR feature selection to handle high dimensional datasets: Vertical partitioning based Iterative MapReduce framework
von: Vivek, Yelleti, et al.
Veröffentlicht: (2022)
von: Vivek, Yelleti, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
von: Gond, Raja, et al.
Veröffentlicht: (2025) -
Towards CXL Resilience to CPU Failures
von: Psistakis, Antonis, et al.
Veröffentlicht: (2026) -
OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
von: Wu, Siyu, et al.
Veröffentlicht: (2025) -
CXL Shared Memory Programming: Barely Distributed and Almost Persistent
von: Xu, Yi, et al.
Veröffentlicht: (2024) -
Modeling the Potential of Message-Free Communication via CXL.mem
von: Vanecek, Stepan, et al.
Veröffentlicht: (2025)