Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Xinjun, Hu, Qingda, Li, Junru, Li, Feifei, Zhu, Yicong, Zhou, Yuqi, Lin, Qiuru, Dai, Jian, Kong, Yang, Zhang, Jiayu, Xu, Guoqiang, Liu, Qiang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PolarStore: High-Performance Data Compression for Large-Scale Cloud-Native Databases
di: Hu, Qingda, et al.
Pubblicazione: (2025)
di: Hu, Qingda, et al.
Pubblicazione: (2025)
CXL Shared Memory Programming: Barely Distributed and Almost Persistent
di: Xu, Yi, et al.
Pubblicazione: (2024)
di: Xu, Yi, et al.
Pubblicazione: (2024)
Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter
di: Qin, Ruoyu, et al.
Pubblicazione: (2026)
di: Qin, Ruoyu, et al.
Pubblicazione: (2026)
Self-Evolving Distributed Memory Architecture for Scalable AI Systems
di: Li, Zixuan, et al.
Pubblicazione: (2026)
di: Li, Zixuan, et al.
Pubblicazione: (2026)
KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud Provider
di: Wang, Jiahao, et al.
Pubblicazione: (2025)
di: Wang, Jiahao, et al.
Pubblicazione: (2025)
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
di: Qin, Ruoyu, et al.
Pubblicazione: (2024)
di: Qin, Ruoyu, et al.
Pubblicazione: (2024)
HybridTier: an Adaptive and Lightweight CXL-Memory Tiering System
di: Song, Kevin, et al.
Pubblicazione: (2023)
di: Song, Kevin, et al.
Pubblicazione: (2023)
Analysis and Optimized CXL-Attached Memory Allocation for Long-Context LLM Fine-Tuning
di: Liaw, Yong-Cheng, et al.
Pubblicazione: (2025)
di: Liaw, Yong-Cheng, et al.
Pubblicazione: (2025)
CCCL: Node-Spanning GPU Collectives with CXL Memory Pooling
di: Xu, Dong, et al.
Pubblicazione: (2026)
di: Xu, Dong, et al.
Pubblicazione: (2026)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
di: Yoon, Dongha, et al.
Pubblicazione: (2025)
di: Yoon, Dongha, et al.
Pubblicazione: (2025)
Towards CXL Resilience to CPU Failures
di: Psistakis, Antonis, et al.
Pubblicazione: (2026)
di: Psistakis, Antonis, et al.
Pubblicazione: (2026)
A Programming Model for Disaggregated Memory over CXL
di: Assa, Gal, et al.
Pubblicazione: (2024)
di: Assa, Gal, et al.
Pubblicazione: (2024)
Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
di: Lin, Bin, et al.
Pubblicazione: (2024)
di: Lin, Bin, et al.
Pubblicazione: (2024)
Telepathic Datacenters: Fast RPCs using Shared CXL Memory
di: Mahar, Suyash, et al.
Pubblicazione: (2024)
di: Mahar, Suyash, et al.
Pubblicazione: (2024)
Equilibria: Fair Multi-Tenant CXL Memory Tiering At Scale
di: Zhao, Kaiyang, et al.
Pubblicazione: (2026)
di: Zhao, Kaiyang, et al.
Pubblicazione: (2026)
Matrix representation and GPU-optimized parallel B-spline computing
di: Wu, Jiayu, et al.
Pubblicazione: (2025)
di: Wu, Jiayu, et al.
Pubblicazione: (2025)
Bridging Cache-Friendliness and Concurrency: A Locality-Optimized In-Memory B-Skiplist
di: Luo, Yicong, et al.
Pubblicazione: (2025)
di: Luo, Yicong, et al.
Pubblicazione: (2025)
Modeling the Potential of Message-Free Communication via CXL.mem
di: Vanecek, Stepan, et al.
Pubblicazione: (2025)
di: Vanecek, Stepan, et al.
Pubblicazione: (2025)
emucxl: an emulation framework for CXL-based disaggregated memory applications
di: Gond, Raja, et al.
Pubblicazione: (2024)
di: Gond, Raja, et al.
Pubblicazione: (2024)
Pooling Engram Conditional Memory in Large Language Models using CXL
di: Ma, Ruiyang, et al.
Pubblicazione: (2026)
di: Ma, Ruiyang, et al.
Pubblicazione: (2026)
MPI-over-CXL: Enhancing Communication Efficiency in Distributed HPC Systems
di: Kwon, Miryeong, et al.
Pubblicazione: (2025)
di: Kwon, Miryeong, et al.
Pubblicazione: (2025)
DRust: Language-Guided Distributed Shared Memory with Fine Granularity, Full Transparency, and Ultra Efficiency
di: Ma, Haoran, et al.
Pubblicazione: (2024)
di: Ma, Haoran, et al.
Pubblicazione: (2024)
DPC: A Distributed Page Cache over CXL
di: Bergman, Shai, et al.
Pubblicazione: (2026)
di: Bergman, Shai, et al.
Pubblicazione: (2026)
Hestia: Hyperthread-Level Scheduling for Cloud Microservices with Interference-Aware Attention
di: Yang, Dingyu, et al.
Pubblicazione: (2026)
di: Yang, Dingyu, et al.
Pubblicazione: (2026)
Automatic BLAS Offloading on Unified Memory Architecture: A Study on NVIDIA Grace-Hopper
di: Li, Junjie, et al.
Pubblicazione: (2024)
di: Li, Junjie, et al.
Pubblicazione: (2024)
FLARE: A Dataflow-Aware and Scalable Hardware Architecture for Neural-Hybrid Scientific Lossy Compression
di: Jia, Wenqi, et al.
Pubblicazione: (2025)
di: Jia, Wenqi, et al.
Pubblicazione: (2025)
ScalePool: Hybrid XLink-CXL Fabric for Composable Resource Disaggregation in Unified Scale-up Domains
di: Woo, Hyein, et al.
Pubblicazione: (2025)
di: Woo, Hyein, et al.
Pubblicazione: (2025)
Exploring and Evaluating Real-world CXL: Use Cases and System Adoption
di: Wang, Xi, et al.
Pubblicazione: (2024)
di: Wang, Xi, et al.
Pubblicazione: (2024)
Beluga: Block Synchronization for BFT Consensus Protocols
di: Kichidis, Tasos, et al.
Pubblicazione: (2025)
di: Kichidis, Tasos, et al.
Pubblicazione: (2025)
Jenga: Effective Memory Management for Serving LLM with Heterogeneity
di: Zhang, Chen, et al.
Pubblicazione: (2025)
di: Zhang, Chen, et al.
Pubblicazione: (2025)
ESS: An Offload-Centric Latent-Cache Management Architecture for DeepSeek-V3.2-Exp
di: Chen, Xinhang, et al.
Pubblicazione: (2025)
di: Chen, Xinhang, et al.
Pubblicazione: (2025)
SwitchDelta: Asynchronous Metadata Updating for Distributed Storage with In-Network Data Visibility
di: Li, Junru, et al.
Pubblicazione: (2025)
di: Li, Junru, et al.
Pubblicazione: (2025)
Scalable Distributed Vector Search via Accuracy Preserving Index Construction
di: Xu, Yuming, et al.
Pubblicazione: (2025)
di: Xu, Yuming, et al.
Pubblicazione: (2025)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
di: Zhang, Zhexiang, et al.
Pubblicazione: (2025)
di: Zhang, Zhexiang, et al.
Pubblicazione: (2025)
W4A16 Mixed-Precision Matrix Multiplication on Decoupled Architecture: Kernel Design and Memory Bottleneck Analysis for Ascend NPUs
di: He, Yuanhong, et al.
Pubblicazione: (2026)
di: He, Yuanhong, et al.
Pubblicazione: (2026)
On the Performance and Memory Footprint of Distributed Training: An Empirical Study on Transformers
di: Lu, Zhengxian, et al.
Pubblicazione: (2024)
di: Lu, Zhengxian, et al.
Pubblicazione: (2024)
Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration
di: Li, Zhonggen, et al.
Pubblicazione: (2025)
di: Li, Zhonggen, et al.
Pubblicazione: (2025)
Memory-aware Adaptive Scheduling of Scientific Workflows on Heterogeneous Architectures
di: Kulagina, Svetlana, et al.
Pubblicazione: (2025)
di: Kulagina, Svetlana, et al.
Pubblicazione: (2025)
Scalable HPC Job Scheduling and Resource Management in SST
di: Abdurahman, Abubeker, et al.
Pubblicazione: (2025)
di: Abdurahman, Abubeker, et al.
Pubblicazione: (2025)
Parallel Data Object Creation: Towards Scalable Metadata Management in High-Performance I/O Library
di: Li, Youjia, et al.
Pubblicazione: (2025)
di: Li, Youjia, et al.
Pubblicazione: (2025)
Documenti analoghi
-
PolarStore: High-Performance Data Compression for Large-Scale Cloud-Native Databases
di: Hu, Qingda, et al.
Pubblicazione: (2025) -
CXL Shared Memory Programming: Barely Distributed and Almost Persistent
di: Xu, Yi, et al.
Pubblicazione: (2024) -
Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter
di: Qin, Ruoyu, et al.
Pubblicazione: (2026) -
Self-Evolving Distributed Memory Architecture for Scalable AI Systems
di: Li, Zixuan, et al.
Pubblicazione: (2026) -
KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud Provider
di: Wang, Jiahao, et al.
Pubblicazione: (2025)