Disaggregated Multi-Tower: Topology-aware Modeling Technique for Efficient Large-Scale Recommendation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Luo, Liang, Zhang, Buyun, Tsang, Michael, Ma, Yinbin, Chu, Ching-Hsiang, Chen, Yuxin, Li, Shen, Hao, Yuchen, Zhao, Yanli, Lakshminarayanan, Guna, Wen, Ellie Dingqiao, Park, Jongsoo, Mudigere, Dheevatsa, Naumov, Maxim |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Wukong: Towards a Scaling Law for Large-Scale Recommendation
von: Zhang, Buyun, et al.
Veröffentlicht: (2024)
von: Zhang, Buyun, et al.
Veröffentlicht: (2024)
Beyond the Buzz: A Pragmatic Take on Inference Disaggregation
von: Mitra, Tiyasa, et al.
Veröffentlicht: (2025)
von: Mitra, Tiyasa, et al.
Veröffentlicht: (2025)
Nonuniform-Tensor-Parallelism: Mitigating GPU failure impact for Scaled-up LLM Training
von: Arfeen, Daiyaan, et al.
Veröffentlicht: (2025)
von: Arfeen, Daiyaan, et al.
Veröffentlicht: (2025)
Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding
von: Bhatia, Nidhi, et al.
Veröffentlicht: (2025)
von: Bhatia, Nidhi, et al.
Veröffentlicht: (2025)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
Lotus: Optimizing Disaggregated Transactions with Disaggregated Locks
von: Hu, Zhisheng, et al.
Veröffentlicht: (2025)
von: Hu, Zhisheng, et al.
Veröffentlicht: (2025)
Accelerating Communication in Deep Learning Recommendation Model Training with Dual-Level Adaptive Lossy Compression
von: Feng, Hao, et al.
Veröffentlicht: (2024)
von: Feng, Hao, et al.
Veröffentlicht: (2024)
Survey of Disaggregated Memory: Cross-layer Technique Insights for Next-Generation Datacenters
von: Wang, Jing, et al.
Veröffentlicht: (2025)
von: Wang, Jing, et al.
Veröffentlicht: (2025)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
Implicit Turn-Wise Policy Optimization for Proactive User-LLM Interaction
von: Wang, Haoyu, et al.
Veröffentlicht: (2026)
von: Wang, Haoyu, et al.
Veröffentlicht: (2026)
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
von: Li, Rongzhi, et al.
Veröffentlicht: (2025)
von: Li, Rongzhi, et al.
Veröffentlicht: (2025)
Software Resource Disaggregation for HPC with Serverless Computing
von: Copik, Marcin, et al.
Veröffentlicht: (2024)
von: Copik, Marcin, et al.
Veröffentlicht: (2024)
DRackSim: Simulator for Rack-scale Memory Disaggregation
von: Puri, Amit, et al.
Veröffentlicht: (2023)
von: Puri, Amit, et al.
Veröffentlicht: (2023)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
von: Kumar, Madabattula Rajesh, et al.
Veröffentlicht: (2025)
von: Kumar, Madabattula Rajesh, et al.
Veröffentlicht: (2025)
Towards Disaggregation-Native Data Streaming between Devices
von: Asmussen, Nils, et al.
Veröffentlicht: (2024)
von: Asmussen, Nils, et al.
Veröffentlicht: (2024)
SWARM: Replicating Shared Disaggregated-Memory Data in No Time
von: Murat, Antoine, et al.
Veröffentlicht: (2024)
von: Murat, Antoine, et al.
Veröffentlicht: (2024)
DecLock: A Case of Decoupled Locking for Disaggregated Memory
von: Zhang, Hanze, et al.
Veröffentlicht: (2025)
von: Zhang, Hanze, et al.
Veröffentlicht: (2025)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
von: Wu, Yu, et al.
Veröffentlicht: (2025)
von: Wu, Yu, et al.
Veröffentlicht: (2025)
Efficient Multi-round LLM Inference over Disaggregated Serving
von: He, Wenhao, et al.
Veröffentlicht: (2026)
von: He, Wenhao, et al.
Veröffentlicht: (2026)
Proceedings of 3rd Workshop on Heterogeneous Composable and Disaggregated Systems
von: Pinto, Christian, et al.
Veröffentlicht: (2024)
von: Pinto, Christian, et al.
Veröffentlicht: (2024)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025)
TASP: Topology-aware Sequence Parallelism
von: Wang, Yida, et al.
Veröffentlicht: (2025)
von: Wang, Yida, et al.
Veröffentlicht: (2025)
Disaggregated Design for GPU-Based Volumetric Data Structures
von: Meneghin, Massimiliano, et al.
Veröffentlicht: (2025)
von: Meneghin, Massimiliano, et al.
Veröffentlicht: (2025)
OffloadFS: Leveraging Disaggregated Storage for Computation Offloading
von: Moon, Sungho, et al.
Veröffentlicht: (2026)
von: Moon, Sungho, et al.
Veröffentlicht: (2026)
RcLLM: Accelerating Generative Recommendation via Beyond-Prefix KV Caching
von: Zhao, Zhan, et al.
Veröffentlicht: (2026)
von: Zhao, Zhan, et al.
Veröffentlicht: (2026)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
von: Wang, Qipeng
Veröffentlicht: (2026)
von: Wang, Qipeng
Veröffentlicht: (2026)
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
PULSE: Accelerating Distributed Pointer-Traversals on Disaggregated Memory (Extended Version)
von: Tang, Yupeng, et al.
Veröffentlicht: (2023)
von: Tang, Yupeng, et al.
Veröffentlicht: (2023)
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
von: Patke, Archit, et al.
Veröffentlicht: (2025)
von: Patke, Archit, et al.
Veröffentlicht: (2025)
Optimising Virtual Resource Mapping in Multi-Level NUMA Disaggregated Systems
von: Lakew, Ewnetu Bayuh, et al.
Veröffentlicht: (2025)
von: Lakew, Ewnetu Bayuh, et al.
Veröffentlicht: (2025)
KVDirect: Distributed Disaggregated LLM Inference
von: Chen, Shiyang, et al.
Veröffentlicht: (2024)
von: Chen, Shiyang, et al.
Veröffentlicht: (2024)
Cache Coherence Over Disaggregated Memory
von: Wang, Ruihong, et al.
Veröffentlicht: (2024)
von: Wang, Ruihong, et al.
Veröffentlicht: (2024)
DEX: Scalable Range Indexing on Disaggregated Memory [Extended Version]
von: Lu, Baotong, et al.
Veröffentlicht: (2024)
von: Lu, Baotong, et al.
Veröffentlicht: (2024)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
von: Lai, Ruiqi, et al.
Veröffentlicht: (2025)
von: Lai, Ruiqi, et al.
Veröffentlicht: (2025)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
von: Wang, Chao, et al.
Veröffentlicht: (2025)
von: Wang, Chao, et al.
Veröffentlicht: (2025)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
von: Dong, Xianzhe, et al.
Veröffentlicht: (2025)
von: Dong, Xianzhe, et al.
Veröffentlicht: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Wukong: Towards a Scaling Law for Large-Scale Recommendation
von: Zhang, Buyun, et al.
Veröffentlicht: (2024) -
Beyond the Buzz: A Pragmatic Take on Inference Disaggregation
von: Mitra, Tiyasa, et al.
Veröffentlicht: (2025) -
Nonuniform-Tensor-Parallelism: Mitigating GPU failure impact for Scaled-up LLM Training
von: Arfeen, Daiyaan, et al.
Veröffentlicht: (2025) -
Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding
von: Bhatia, Nidhi, et al.
Veröffentlicht: (2025) -
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)