Towards Disaggregation-Native Data Streaming between Devices
Fuente:
arXiv
Salvato in:
| Autori principali: | Asmussen, Nils, Roitzsch, Michael |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Lotus: Optimizing Disaggregated Transactions with Disaggregated Locks
di: Hu, Zhisheng, et al.
Pubblicazione: (2025)
di: Hu, Zhisheng, et al.
Pubblicazione: (2025)
SWARM: Replicating Shared Disaggregated-Memory Data in No Time
di: Murat, Antoine, et al.
Pubblicazione: (2024)
di: Murat, Antoine, et al.
Pubblicazione: (2024)
StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
di: Zhong, Yinmin, et al.
Pubblicazione: (2025)
di: Zhong, Yinmin, et al.
Pubblicazione: (2025)
P/D-Device: Disaggregated Large Language Model between Cloud and Devices
di: Jin, Yibo, et al.
Pubblicazione: (2025)
di: Jin, Yibo, et al.
Pubblicazione: (2025)
DOLMA: A Data Object Level Memory Disaggregation Framework for HPC Applications
di: Zheng, Haoyu, et al.
Pubblicazione: (2025)
di: Zheng, Haoyu, et al.
Pubblicazione: (2025)
Disaggregated Design for GPU-Based Volumetric Data Structures
di: Meneghin, Massimiliano, et al.
Pubblicazione: (2025)
di: Meneghin, Massimiliano, et al.
Pubblicazione: (2025)
Software Resource Disaggregation for HPC with Serverless Computing
di: Copik, Marcin, et al.
Pubblicazione: (2024)
di: Copik, Marcin, et al.
Pubblicazione: (2024)
DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
di: Zhang, Zili, et al.
Pubblicazione: (2024)
di: Zhang, Zili, et al.
Pubblicazione: (2024)
TENT: A Declarative Slice Spraying Engine for Performant and Resilient Data Movement in Disaggregated LLM Serving
di: Ren, Feng, et al.
Pubblicazione: (2026)
di: Ren, Feng, et al.
Pubblicazione: (2026)
DRackSim: Simulator for Rack-scale Memory Disaggregation
di: Puri, Amit, et al.
Pubblicazione: (2023)
di: Puri, Amit, et al.
Pubblicazione: (2023)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
di: Kumar, Madabattula Rajesh, et al.
Pubblicazione: (2025)
di: Kumar, Madabattula Rajesh, et al.
Pubblicazione: (2025)
Towards an Adaptive Runtime System for Cloud-Native HPC
di: Bhosale, Aditya, et al.
Pubblicazione: (2026)
di: Bhosale, Aditya, et al.
Pubblicazione: (2026)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
di: Kumar, Satyam, et al.
Pubblicazione: (2026)
di: Kumar, Satyam, et al.
Pubblicazione: (2026)
Proceedings of 3rd Workshop on Heterogeneous Composable and Disaggregated Systems
di: Pinto, Christian, et al.
Pubblicazione: (2024)
di: Pinto, Christian, et al.
Pubblicazione: (2024)
DecLock: A Case of Decoupled Locking for Disaggregated Memory
di: Zhang, Hanze, et al.
Pubblicazione: (2025)
di: Zhang, Hanze, et al.
Pubblicazione: (2025)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
di: Wu, Yu, et al.
Pubblicazione: (2025)
di: Wu, Yu, et al.
Pubblicazione: (2025)
Efficient Multi-round LLM Inference over Disaggregated Serving
di: He, Wenhao, et al.
Pubblicazione: (2026)
di: He, Wenhao, et al.
Pubblicazione: (2026)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
di: Zhang, Zhexiang, et al.
Pubblicazione: (2025)
di: Zhang, Zhexiang, et al.
Pubblicazione: (2025)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
di: Wang, Qipeng
Pubblicazione: (2026)
di: Wang, Qipeng
Pubblicazione: (2026)
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
di: Wu, Tianyuan, et al.
Pubblicazione: (2025)
di: Wu, Tianyuan, et al.
Pubblicazione: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
PULSE: Accelerating Distributed Pointer-Traversals on Disaggregated Memory (Extended Version)
di: Tang, Yupeng, et al.
Pubblicazione: (2023)
di: Tang, Yupeng, et al.
Pubblicazione: (2023)
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
di: Patke, Archit, et al.
Pubblicazione: (2025)
di: Patke, Archit, et al.
Pubblicazione: (2025)
Optimising Virtual Resource Mapping in Multi-Level NUMA Disaggregated Systems
di: Lakew, Ewnetu Bayuh, et al.
Pubblicazione: (2025)
di: Lakew, Ewnetu Bayuh, et al.
Pubblicazione: (2025)
Low Latency, High Bandwidth Streaming of Experimental Data with EJFAT
di: Baldin, Ilya, et al.
Pubblicazione: (2025)
di: Baldin, Ilya, et al.
Pubblicazione: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
Disaggregated Memory with SmartNIC Offloading: a Case Study on Graph Processing
di: Wahlgren, Jacob, et al.
Pubblicazione: (2024)
di: Wahlgren, Jacob, et al.
Pubblicazione: (2024)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
di: Lai, Ruiqi, et al.
Pubblicazione: (2025)
di: Lai, Ruiqi, et al.
Pubblicazione: (2025)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
di: Wang, Chao, et al.
Pubblicazione: (2025)
di: Wang, Chao, et al.
Pubblicazione: (2025)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
di: Dong, Xianzhe, et al.
Pubblicazione: (2025)
di: Dong, Xianzhe, et al.
Pubblicazione: (2025)
FlexKV: Flexible Index Offloading for Memory-Disaggregated Key-Value Store
di: Hu, Zhisheng, et al.
Pubblicazione: (2025)
di: Hu, Zhisheng, et al.
Pubblicazione: (2025)
uBFT: Microsecond-scale BFT using Disaggregated Memory [Extended Version]
di: Aguilera, Marcos K., et al.
Pubblicazione: (2022)
di: Aguilera, Marcos K., et al.
Pubblicazione: (2022)
vLLM-Omni: Fully Disaggregated Serving for Any-to-Any Multimodal Models
di: Yin, Peiqi, et al.
Pubblicazione: (2026)
di: Yin, Peiqi, et al.
Pubblicazione: (2026)
DiFache: Efficient and Scalable Caching on Disaggregated Memory using Decentralized Coherence
di: Zhang, Hanze, et al.
Pubblicazione: (2025)
di: Zhang, Hanze, et al.
Pubblicazione: (2025)
A Study on Messaging Trade-offs in Data Streaming for Scientific Workflows
di: George, Anjus, et al.
Pubblicazione: (2025)
di: George, Anjus, et al.
Pubblicazione: (2025)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
di: Zhong, Yinmin, et al.
Pubblicazione: (2024)
di: Zhong, Yinmin, et al.
Pubblicazione: (2024)
TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
di: Zhang, Hongbin, et al.
Pubblicazione: (2025)
di: Zhang, Hongbin, et al.
Pubblicazione: (2025)
Revealing the Challenges of Attention-FFN Disaggregation for Modern MoE Models and Hardware Systems
di: Liu, Guowei, et al.
Pubblicazione: (2026)
di: Liu, Guowei, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Lotus: Optimizing Disaggregated Transactions with Disaggregated Locks
di: Hu, Zhisheng, et al.
Pubblicazione: (2025) -
SWARM: Replicating Shared Disaggregated-Memory Data in No Time
di: Murat, Antoine, et al.
Pubblicazione: (2024) -
StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
di: Zhong, Yinmin, et al.
Pubblicazione: (2025) -
P/D-Device: Disaggregated Large Language Model between Cloud and Devices
di: Jin, Yibo, et al.
Pubblicazione: (2025) -
DOLMA: A Data Object Level Memory Disaggregation Framework for HPC Applications
di: Zheng, Haoyu, et al.
Pubblicazione: (2025)