DiFache: Efficient and Scalable Caching on Disaggregated Memory using Decentralized Coherence
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Hanze, Wang, Kaiming, Chen, Rong, Wei, Xingda, Chen, Haibo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DecLock: A Case of Decoupled Locking for Disaggregated Memory
by: Zhang, Hanze, et al.
Published: (2025)
by: Zhang, Hanze, et al.
Published: (2025)
Towards Lock Modularization for Heterogeneous Environments
by: Zhang, Hanze, et al.
Published: (2025)
by: Zhang, Hanze, et al.
Published: (2025)
KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
by: Cheng, Rongxin, et al.
Published: (2024)
by: Cheng, Rongxin, et al.
Published: (2024)
Cache Coherence Over Disaggregated Memory
by: Wang, Ruihong, et al.
Published: (2024)
by: Wang, Ruihong, et al.
Published: (2024)
BLITZSCALE: Fast and Live Large Model Autoscaling with O(1) Host Caching
by: Zhang, Dingyan, et al.
Published: (2024)
by: Zhang, Dingyan, et al.
Published: (2024)
KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud Provider
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
by: Zhang, Zhexiang, et al.
Published: (2025)
by: Zhang, Zhexiang, et al.
Published: (2025)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
by: Yoon, Dongha, et al.
Published: (2025)
by: Yoon, Dongha, et al.
Published: (2025)
DEX: Scalable Range Indexing on Disaggregated Memory [Extended Version]
by: Lu, Baotong, et al.
Published: (2024)
by: Lu, Baotong, et al.
Published: (2024)
Characterizing the Dilemma of Performance and Index Size in Billion-Scale Vector Search and Breaking It with Second-Tier Memory
by: Cheng, Rongxin, et al.
Published: (2024)
by: Cheng, Rongxin, et al.
Published: (2024)
ProMoE: Fast MoE-based LLM Serving using Proactive Caching
by: Song, Xiaoniu, et al.
Published: (2024)
by: Song, Xiaoniu, et al.
Published: (2024)
Lotus: Optimizing Disaggregated Transactions with Disaggregated Locks
by: Hu, Zhisheng, et al.
Published: (2025)
by: Hu, Zhisheng, et al.
Published: (2025)
TDC-Cache: A Trustworthy Decentralized Cooperative Caching Framework for Web3.0
by: Chen, Jinyu, et al.
Published: (2025)
by: Chen, Jinyu, et al.
Published: (2025)
uBFT: Microsecond-scale BFT using Disaggregated Memory [Extended Version]
by: Aguilera, Marcos K., et al.
Published: (2022)
by: Aguilera, Marcos K., et al.
Published: (2022)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
by: Chen, Xing, et al.
Published: (2025)
by: Chen, Xing, et al.
Published: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
by: He, Yiyuan, et al.
Published: (2025)
by: He, Yiyuan, et al.
Published: (2025)
DRackSim: Simulator for Rack-scale Memory Disaggregation
by: Puri, Amit, et al.
Published: (2023)
by: Puri, Amit, et al.
Published: (2023)
SWARM: Replicating Shared Disaggregated-Memory Data in No Time
by: Murat, Antoine, et al.
Published: (2024)
by: Murat, Antoine, et al.
Published: (2024)
PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated Speculation
by: Wei, Xingda, et al.
Published: (2024)
by: Wei, Xingda, et al.
Published: (2024)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
by: Qianli, Liu, et al.
Published: (2025)
by: Qianli, Liu, et al.
Published: (2025)
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
by: Patke, Archit, et al.
Published: (2025)
by: Patke, Archit, et al.
Published: (2025)
PULSE: Accelerating Distributed Pointer-Traversals on Disaggregated Memory (Extended Version)
by: Tang, Yupeng, et al.
Published: (2023)
by: Tang, Yupeng, et al.
Published: (2023)
TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
by: Zhang, Hongbin, et al.
Published: (2025)
by: Zhang, Hongbin, et al.
Published: (2025)
The Dawn of Disaggregation and the Coherence Conundrum: A Call for Federated Coherence
by: Hong, Jaewan, et al.
Published: (2025)
by: Hong, Jaewan, et al.
Published: (2025)
DOLMA: A Data Object Level Memory Disaggregation Framework for HPC Applications
by: Zheng, Haoyu, et al.
Published: (2025)
by: Zheng, Haoyu, et al.
Published: (2025)
FlexKV: Flexible Index Offloading for Memory-Disaggregated Key-Value Store
by: Hu, Zhisheng, et al.
Published: (2025)
by: Hu, Zhisheng, et al.
Published: (2025)
Disaggregated Memory with SmartNIC Offloading: a Case Study on Graph Processing
by: Wahlgren, Jacob, et al.
Published: (2024)
by: Wahlgren, Jacob, et al.
Published: (2024)
DiT-HC: Enabling Efficient Training of Visual Generation Model DiT on HPC-oriented CPU Cluster
by: Zhang, Jinxiao, et al.
Published: (2026)
by: Zhang, Jinxiao, et al.
Published: (2026)
Efficient Multi-round LLM Inference over Disaggregated Serving
by: He, Wenhao, et al.
Published: (2026)
by: He, Wenhao, et al.
Published: (2026)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
CIDER: Boosting Memory-Disaggregated Key-Value Stores with Pessimistic Synchronization
by: Du, Yuxuan, et al.
Published: (2026)
by: Du, Yuxuan, et al.
Published: (2026)
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
by: Wu, Tianyuan, et al.
Published: (2025)
by: Wu, Tianyuan, et al.
Published: (2025)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
by: Wang, Chao, et al.
Published: (2025)
by: Wang, Chao, et al.
Published: (2025)
StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
by: Zhong, Yinmin, et al.
Published: (2025)
by: Zhong, Yinmin, et al.
Published: (2025)
Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation
by: Chen, Shaoyuan, et al.
Published: (2024)
by: Chen, Shaoyuan, et al.
Published: (2024)
Towards Efficient and Scalable Distributed Vector Search with RDMA
by: Zhi, Xiangyu, et al.
Published: (2025)
by: Zhi, Xiangyu, et al.
Published: (2025)
HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
by: Ruan, Chaoyi, et al.
Published: (2025)
by: Ruan, Chaoyi, et al.
Published: (2025)
CacheFL: Privacy-Preserving and Efficient Federated Cache Model Fine-Tuning for Vision-Language Models
by: Yi, Mengjun, et al.
Published: (2025)
by: Yi, Mengjun, et al.
Published: (2025)
Similar Items
-
DecLock: A Case of Decoupled Locking for Disaggregated Memory
by: Zhang, Hanze, et al.
Published: (2025) -
Towards Lock Modularization for Heterogeneous Environments
by: Zhang, Hanze, et al.
Published: (2025) -
KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
by: Cheng, Rongxin, et al.
Published: (2024) -
Cache Coherence Over Disaggregated Memory
by: Wang, Ruihong, et al.
Published: (2024) -
BLITZSCALE: Fast and Live Large Model Autoscaling with O(1) Host Caching
by: Zhang, Dingyan, et al.
Published: (2024)