Towards Disaggregation-Native Data Streaming between Devices
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Asmussen, Nils, Roitzsch, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Lotus: Optimizing Disaggregated Transactions with Disaggregated Locks
von: Hu, Zhisheng, et al.
Veröffentlicht: (2025)
von: Hu, Zhisheng, et al.
Veröffentlicht: (2025)
SWARM: Replicating Shared Disaggregated-Memory Data in No Time
von: Murat, Antoine, et al.
Veröffentlicht: (2024)
von: Murat, Antoine, et al.
Veröffentlicht: (2024)
StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
von: Zhong, Yinmin, et al.
Veröffentlicht: (2025)
von: Zhong, Yinmin, et al.
Veröffentlicht: (2025)
P/D-Device: Disaggregated Large Language Model between Cloud and Devices
von: Jin, Yibo, et al.
Veröffentlicht: (2025)
von: Jin, Yibo, et al.
Veröffentlicht: (2025)
DOLMA: A Data Object Level Memory Disaggregation Framework for HPC Applications
von: Zheng, Haoyu, et al.
Veröffentlicht: (2025)
von: Zheng, Haoyu, et al.
Veröffentlicht: (2025)
Disaggregated Design for GPU-Based Volumetric Data Structures
von: Meneghin, Massimiliano, et al.
Veröffentlicht: (2025)
von: Meneghin, Massimiliano, et al.
Veröffentlicht: (2025)
Software Resource Disaggregation for HPC with Serverless Computing
von: Copik, Marcin, et al.
Veröffentlicht: (2024)
von: Copik, Marcin, et al.
Veröffentlicht: (2024)
DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
von: Zhang, Zili, et al.
Veröffentlicht: (2024)
von: Zhang, Zili, et al.
Veröffentlicht: (2024)
TENT: A Declarative Slice Spraying Engine for Performant and Resilient Data Movement in Disaggregated LLM Serving
von: Ren, Feng, et al.
Veröffentlicht: (2026)
von: Ren, Feng, et al.
Veröffentlicht: (2026)
DRackSim: Simulator for Rack-scale Memory Disaggregation
von: Puri, Amit, et al.
Veröffentlicht: (2023)
von: Puri, Amit, et al.
Veröffentlicht: (2023)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
von: Kumar, Madabattula Rajesh, et al.
Veröffentlicht: (2025)
von: Kumar, Madabattula Rajesh, et al.
Veröffentlicht: (2025)
Towards an Adaptive Runtime System for Cloud-Native HPC
von: Bhosale, Aditya, et al.
Veröffentlicht: (2026)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2026)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
Proceedings of 3rd Workshop on Heterogeneous Composable and Disaggregated Systems
von: Pinto, Christian, et al.
Veröffentlicht: (2024)
von: Pinto, Christian, et al.
Veröffentlicht: (2024)
DecLock: A Case of Decoupled Locking for Disaggregated Memory
von: Zhang, Hanze, et al.
Veröffentlicht: (2025)
von: Zhang, Hanze, et al.
Veröffentlicht: (2025)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
von: Wu, Yu, et al.
Veröffentlicht: (2025)
von: Wu, Yu, et al.
Veröffentlicht: (2025)
Efficient Multi-round LLM Inference over Disaggregated Serving
von: He, Wenhao, et al.
Veröffentlicht: (2026)
von: He, Wenhao, et al.
Veröffentlicht: (2026)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
von: Wang, Qipeng
Veröffentlicht: (2026)
von: Wang, Qipeng
Veröffentlicht: (2026)
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
PULSE: Accelerating Distributed Pointer-Traversals on Disaggregated Memory (Extended Version)
von: Tang, Yupeng, et al.
Veröffentlicht: (2023)
von: Tang, Yupeng, et al.
Veröffentlicht: (2023)
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
von: Patke, Archit, et al.
Veröffentlicht: (2025)
von: Patke, Archit, et al.
Veröffentlicht: (2025)
Optimising Virtual Resource Mapping in Multi-Level NUMA Disaggregated Systems
von: Lakew, Ewnetu Bayuh, et al.
Veröffentlicht: (2025)
von: Lakew, Ewnetu Bayuh, et al.
Veröffentlicht: (2025)
Low Latency, High Bandwidth Streaming of Experimental Data with EJFAT
von: Baldin, Ilya, et al.
Veröffentlicht: (2025)
von: Baldin, Ilya, et al.
Veröffentlicht: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
Disaggregated Memory with SmartNIC Offloading: a Case Study on Graph Processing
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2024)
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2024)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
von: Lai, Ruiqi, et al.
Veröffentlicht: (2025)
von: Lai, Ruiqi, et al.
Veröffentlicht: (2025)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
von: Wang, Chao, et al.
Veröffentlicht: (2025)
von: Wang, Chao, et al.
Veröffentlicht: (2025)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
von: Dong, Xianzhe, et al.
Veröffentlicht: (2025)
von: Dong, Xianzhe, et al.
Veröffentlicht: (2025)
FlexKV: Flexible Index Offloading for Memory-Disaggregated Key-Value Store
von: Hu, Zhisheng, et al.
Veröffentlicht: (2025)
von: Hu, Zhisheng, et al.
Veröffentlicht: (2025)
uBFT: Microsecond-scale BFT using Disaggregated Memory [Extended Version]
von: Aguilera, Marcos K., et al.
Veröffentlicht: (2022)
von: Aguilera, Marcos K., et al.
Veröffentlicht: (2022)
vLLM-Omni: Fully Disaggregated Serving for Any-to-Any Multimodal Models
von: Yin, Peiqi, et al.
Veröffentlicht: (2026)
von: Yin, Peiqi, et al.
Veröffentlicht: (2026)
DiFache: Efficient and Scalable Caching on Disaggregated Memory using Decentralized Coherence
von: Zhang, Hanze, et al.
Veröffentlicht: (2025)
von: Zhang, Hanze, et al.
Veröffentlicht: (2025)
A Study on Messaging Trade-offs in Data Streaming for Scientific Workflows
von: George, Anjus, et al.
Veröffentlicht: (2025)
von: George, Anjus, et al.
Veröffentlicht: (2025)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
von: Zhong, Yinmin, et al.
Veröffentlicht: (2024)
von: Zhong, Yinmin, et al.
Veröffentlicht: (2024)
TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
von: Zhang, Hongbin, et al.
Veröffentlicht: (2025)
von: Zhang, Hongbin, et al.
Veröffentlicht: (2025)
Revealing the Challenges of Attention-FFN Disaggregation for Modern MoE Models and Hardware Systems
von: Liu, Guowei, et al.
Veröffentlicht: (2026)
von: Liu, Guowei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Lotus: Optimizing Disaggregated Transactions with Disaggregated Locks
von: Hu, Zhisheng, et al.
Veröffentlicht: (2025) -
SWARM: Replicating Shared Disaggregated-Memory Data in No Time
von: Murat, Antoine, et al.
Veröffentlicht: (2024) -
StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
von: Zhong, Yinmin, et al.
Veröffentlicht: (2025) -
P/D-Device: Disaggregated Large Language Model between Cloud and Devices
von: Jin, Yibo, et al.
Veröffentlicht: (2025) -
DOLMA: A Data Object Level Memory Disaggregation Framework for HPC Applications
von: Zheng, Haoyu, et al.
Veröffentlicht: (2025)