Lotus: Optimizing Disaggregated Transactions with Disaggregated Locks
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hu, Zhisheng, Zuo, Pengfei, Hu, Junliang, Chen, Yizou, Wang, Yingjia, Yang, Ming-Chang |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
FlexKV: Flexible Index Offloading for Memory-Disaggregated Key-Value Store
par: Hu, Zhisheng, et autres
Publié: (2025)
par: Hu, Zhisheng, et autres
Publié: (2025)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
par: Wang, Chao, et autres
Publié: (2025)
par: Wang, Chao, et autres
Publié: (2025)
DecLock: A Case of Decoupled Locking for Disaggregated Memory
par: Zhang, Hanze, et autres
Publié: (2025)
par: Zhang, Hanze, et autres
Publié: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
par: Hu, Cunchen, et autres
Publié: (2024)
par: Hu, Cunchen, et autres
Publié: (2024)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
par: Hu, Cunchen, et autres
Publié: (2024)
par: Hu, Cunchen, et autres
Publié: (2024)
Injecting Adrenaline into LLM Serving: Boosting Resource Utilization and Throughput via Attention Disaggregation
par: Liang, Yunkai, et autres
Publié: (2025)
par: Liang, Yunkai, et autres
Publié: (2025)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
par: Zhang, Zhexiang, et autres
Publié: (2025)
par: Zhang, Zhexiang, et autres
Publié: (2025)
Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
par: Liu, Yunzhao, et autres
Publié: (2025)
par: Liu, Yunzhao, et autres
Publié: (2025)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
par: Zhong, Yinmin, et autres
Publié: (2024)
par: Zhong, Yinmin, et autres
Publié: (2024)
Software Resource Disaggregation for HPC with Serverless Computing
par: Copik, Marcin, et autres
Publié: (2024)
par: Copik, Marcin, et autres
Publié: (2024)
Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation
par: Hu, Tiancheng, et autres
Publié: (2026)
par: Hu, Tiancheng, et autres
Publié: (2026)
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
par: Basit, Omar, et autres
Publié: (2026)
par: Basit, Omar, et autres
Publié: (2026)
DiFache: Efficient and Scalable Caching on Disaggregated Memory using Decentralized Coherence
par: Zhang, Hanze, et autres
Publié: (2025)
par: Zhang, Hanze, et autres
Publié: (2025)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
par: Wu, Yu, et autres
Publié: (2025)
par: Wu, Yu, et autres
Publié: (2025)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
par: Kumar, Madabattula Rajesh, et autres
Publié: (2025)
par: Kumar, Madabattula Rajesh, et autres
Publié: (2025)
DRackSim: Simulator for Rack-scale Memory Disaggregation
par: Puri, Amit, et autres
Publié: (2023)
par: Puri, Amit, et autres
Publié: (2023)
Towards Disaggregation-Native Data Streaming between Devices
par: Asmussen, Nils, et autres
Publié: (2024)
par: Asmussen, Nils, et autres
Publié: (2024)
SWARM: Replicating Shared Disaggregated-Memory Data in No Time
par: Murat, Antoine, et autres
Publié: (2024)
par: Murat, Antoine, et autres
Publié: (2024)
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
par: Wu, Tianyuan, et autres
Publié: (2025)
par: Wu, Tianyuan, et autres
Publié: (2025)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
par: Wang, Qipeng
Publié: (2026)
par: Wang, Qipeng
Publié: (2026)
vLLM-Omni: Fully Disaggregated Serving for Any-to-Any Multimodal Models
par: Yin, Peiqi, et autres
Publié: (2026)
par: Yin, Peiqi, et autres
Publié: (2026)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
par: Ruan, Chaoyi, et autres
Publié: (2025)
par: Ruan, Chaoyi, et autres
Publié: (2025)
DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
par: Zhang, Zili, et autres
Publié: (2024)
par: Zhang, Zili, et autres
Publié: (2024)
Efficient Multi-round LLM Inference over Disaggregated Serving
par: He, Wenhao, et autres
Publié: (2026)
par: He, Wenhao, et autres
Publié: (2026)
Proceedings of 3rd Workshop on Heterogeneous Composable and Disaggregated Systems
par: Pinto, Christian, et autres
Publié: (2024)
par: Pinto, Christian, et autres
Publié: (2024)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
par: Dong, Xianzhe, et autres
Publié: (2025)
par: Dong, Xianzhe, et autres
Publié: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
par: He, Yiyuan, et autres
Publié: (2025)
par: He, Yiyuan, et autres
Publié: (2025)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
par: Chen, Xing, et autres
Publié: (2025)
par: Chen, Xing, et autres
Publié: (2025)
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
par: Jiang, Youhe, et autres
Publié: (2025)
par: Jiang, Youhe, et autres
Publié: (2025)
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
par: Patke, Archit, et autres
Publié: (2025)
par: Patke, Archit, et autres
Publié: (2025)
Optimising Virtual Resource Mapping in Multi-Level NUMA Disaggregated Systems
par: Lakew, Ewnetu Bayuh, et autres
Publié: (2025)
par: Lakew, Ewnetu Bayuh, et autres
Publié: (2025)
PULSE: Accelerating Distributed Pointer-Traversals on Disaggregated Memory (Extended Version)
par: Tang, Yupeng, et autres
Publié: (2023)
par: Tang, Yupeng, et autres
Publié: (2023)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
par: Chen, Hongyu, et autres
Publié: (2026)
par: Chen, Hongyu, et autres
Publié: (2026)
TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
par: Zhang, Hongbin, et autres
Publié: (2025)
par: Zhang, Hongbin, et autres
Publié: (2025)
DEX: Scalable Range Indexing on Disaggregated Memory [Extended Version]
par: Lu, Baotong, et autres
Publié: (2024)
par: Lu, Baotong, et autres
Publié: (2024)
Disaggregated Design for GPU-Based Volumetric Data Structures
par: Meneghin, Massimiliano, et autres
Publié: (2025)
par: Meneghin, Massimiliano, et autres
Publié: (2025)
OffloadFS: Leveraging Disaggregated Storage for Computation Offloading
par: Moon, Sungho, et autres
Publié: (2026)
par: Moon, Sungho, et autres
Publié: (2026)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
par: Liao, Junhan, et autres
Publié: (2025)
par: Liao, Junhan, et autres
Publié: (2025)
Revealing the Challenges of Attention-FFN Disaggregation for Modern MoE Models and Hardware Systems
par: Liu, Guowei, et autres
Publié: (2026)
par: Liu, Guowei, et autres
Publié: (2026)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
par: Lai, Ruiqi, et autres
Publié: (2025)
par: Lai, Ruiqi, et autres
Publié: (2025)
Documents similaires
-
FlexKV: Flexible Index Offloading for Memory-Disaggregated Key-Value Store
par: Hu, Zhisheng, et autres
Publié: (2025) -
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
par: Wang, Chao, et autres
Publié: (2025) -
DecLock: A Case of Decoupled Locking for Disaggregated Memory
par: Zhang, Hanze, et autres
Publié: (2025) -
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
par: Hu, Cunchen, et autres
Publié: (2024) -
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
par: Hu, Cunchen, et autres
Publié: (2024)