Optimising Virtual Resource Mapping in Multi-Level NUMA Disaggregated Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Lakew, Ewnetu Bayuh, Svärd, Petter, Elmroth, Erik, Tordsson, Johan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A HPC Co-Scheduler with Reinforcement Learning
di: Souza, Abel, et al.
Pubblicazione: (2024)
di: Souza, Abel, et al.
Pubblicazione: (2024)
Silent Failures in Stateless Systems: Rethinking Anomaly Detection for Serverless Computing
di: Nguyen, Chanh, et al.
Pubblicazione: (2025)
di: Nguyen, Chanh, et al.
Pubblicazione: (2025)
Workload Buoyancy: Keeping Apps Afloat by Identifying Shared Resource Bottlenecks
di: Larsson, Oliver, et al.
Pubblicazione: (2026)
di: Larsson, Oliver, et al.
Pubblicazione: (2026)
LA-IMR: Latency-Aware, Predictive In-Memory Routing and Proactive Autoscaling for Tail-Latency-Sensitive Cloud Robotics
di: Seo, Eunil, et al.
Pubblicazione: (2025)
di: Seo, Eunil, et al.
Pubblicazione: (2025)
ASA -- The Adaptive Scheduling Algorithm
di: Souza, Abel, et al.
Pubblicazione: (2024)
di: Souza, Abel, et al.
Pubblicazione: (2024)
Taming Cold Starts: Proactive Serverless Scheduling with Model Predictive Control
di: Nguyen, Chanh, et al.
Pubblicazione: (2025)
di: Nguyen, Chanh, et al.
Pubblicazione: (2025)
Software Resource Disaggregation for HPC with Serverless Computing
di: Copik, Marcin, et al.
Pubblicazione: (2024)
di: Copik, Marcin, et al.
Pubblicazione: (2024)
Speculative Policy Orchestration: A Latency-Resilient Framework for Cloud-Robotic Manipulation
di: Nguyen, Chanh, et al.
Pubblicazione: (2026)
di: Nguyen, Chanh, et al.
Pubblicazione: (2026)
Artifact Evaluation for Distributed Systems: Current Practices and Beyond
di: Sedghpour, Mohammad Reza Saleh, et al.
Pubblicazione: (2024)
di: Sedghpour, Mohammad Reza Saleh, et al.
Pubblicazione: (2024)
Lotus: Optimizing Disaggregated Transactions with Disaggregated Locks
di: Hu, Zhisheng, et al.
Pubblicazione: (2025)
di: Hu, Zhisheng, et al.
Pubblicazione: (2025)
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
di: Wu, Tianyuan, et al.
Pubblicazione: (2025)
di: Wu, Tianyuan, et al.
Pubblicazione: (2025)
DOLMA: A Data Object Level Memory Disaggregation Framework for HPC Applications
di: Zheng, Haoyu, et al.
Pubblicazione: (2025)
di: Zheng, Haoyu, et al.
Pubblicazione: (2025)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
di: Chen, Xing, et al.
Pubblicazione: (2025)
di: Chen, Xing, et al.
Pubblicazione: (2025)
Trust-Aware Routing for Distributed Generative AI Inference at the Edge
di: Nguyen, Chanh, et al.
Pubblicazione: (2026)
di: Nguyen, Chanh, et al.
Pubblicazione: (2026)
Efficient Multi-round LLM Inference over Disaggregated Serving
di: He, Wenhao, et al.
Pubblicazione: (2026)
di: He, Wenhao, et al.
Pubblicazione: (2026)
Proceedings of 3rd Workshop on Heterogeneous Composable and Disaggregated Systems
di: Pinto, Christian, et al.
Pubblicazione: (2024)
di: Pinto, Christian, et al.
Pubblicazione: (2024)
ScalePool: Hybrid XLink-CXL Fabric for Composable Resource Disaggregation in Unified Scale-up Domains
di: Woo, Hyein, et al.
Pubblicazione: (2025)
di: Woo, Hyein, et al.
Pubblicazione: (2025)
Hardware-Level QoS Enforcement Features: Technologies, Use Cases, and Research Challenges
di: Larsson, Oliver, et al.
Pubblicazione: (2025)
di: Larsson, Oliver, et al.
Pubblicazione: (2025)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
di: Chen, Hongyu, et al.
Pubblicazione: (2026)
di: Chen, Hongyu, et al.
Pubblicazione: (2026)
Revealing the Challenges of Attention-FFN Disaggregation for Modern MoE Models and Hardware Systems
di: Liu, Guowei, et al.
Pubblicazione: (2026)
di: Liu, Guowei, et al.
Pubblicazione: (2026)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
di: Kumar, Madabattula Rajesh, et al.
Pubblicazione: (2025)
di: Kumar, Madabattula Rajesh, et al.
Pubblicazione: (2025)
DRackSim: Simulator for Rack-scale Memory Disaggregation
di: Puri, Amit, et al.
Pubblicazione: (2023)
di: Puri, Amit, et al.
Pubblicazione: (2023)
Towards Disaggregation-Native Data Streaming between Devices
di: Asmussen, Nils, et al.
Pubblicazione: (2024)
di: Asmussen, Nils, et al.
Pubblicazione: (2024)
SWARM: Replicating Shared Disaggregated-Memory Data in No Time
di: Murat, Antoine, et al.
Pubblicazione: (2024)
di: Murat, Antoine, et al.
Pubblicazione: (2024)
Enabling Disaggregated Multi-Stage MLLM Inference via GPU-Internal Scheduling and Resource Sharing
di: Zhao, Lingxiao, et al.
Pubblicazione: (2025)
di: Zhao, Lingxiao, et al.
Pubblicazione: (2025)
DecLock: A Case of Decoupled Locking for Disaggregated Memory
di: Zhang, Hanze, et al.
Pubblicazione: (2025)
di: Zhang, Hanze, et al.
Pubblicazione: (2025)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
di: Wu, Yu, et al.
Pubblicazione: (2025)
di: Wu, Yu, et al.
Pubblicazione: (2025)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
di: Zhang, Zhexiang, et al.
Pubblicazione: (2025)
di: Zhang, Zhexiang, et al.
Pubblicazione: (2025)
SmartPQ: An Adaptive Concurrent Priority Queue for NUMA Architectures
di: Giannoula, Christina, et al.
Pubblicazione: (2024)
di: Giannoula, Christina, et al.
Pubblicazione: (2024)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
di: Patke, Archit, et al.
Pubblicazione: (2025)
di: Patke, Archit, et al.
Pubblicazione: (2025)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
di: Wang, Qipeng
Pubblicazione: (2026)
di: Wang, Qipeng
Pubblicazione: (2026)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
PULSE: Accelerating Distributed Pointer-Traversals on Disaggregated Memory (Extended Version)
di: Tang, Yupeng, et al.
Pubblicazione: (2023)
di: Tang, Yupeng, et al.
Pubblicazione: (2023)
Mao: Machine learning approach for NUMA optimization in Warehouse Scale Computers
di: Liu, Yueji, et al.
Pubblicazione: (2024)
di: Liu, Yueji, et al.
Pubblicazione: (2024)
An Incremental Multi-Level, Multi-Scale Approach to Assessment of Multifidelity HPC Systems
di: Shilpika, Shilpika, et al.
Pubblicazione: (2025)
di: Shilpika, Shilpika, et al.
Pubblicazione: (2025)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
di: Lai, Ruiqi, et al.
Pubblicazione: (2025)
di: Lai, Ruiqi, et al.
Pubblicazione: (2025)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
di: Wang, Chao, et al.
Pubblicazione: (2025)
di: Wang, Chao, et al.
Pubblicazione: (2025)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
di: Dong, Xianzhe, et al.
Pubblicazione: (2025)
di: Dong, Xianzhe, et al.
Pubblicazione: (2025)
Documenti analoghi
-
A HPC Co-Scheduler with Reinforcement Learning
di: Souza, Abel, et al.
Pubblicazione: (2024) -
Silent Failures in Stateless Systems: Rethinking Anomaly Detection for Serverless Computing
di: Nguyen, Chanh, et al.
Pubblicazione: (2025) -
Workload Buoyancy: Keeping Apps Afloat by Identifying Shared Resource Bottlenecks
di: Larsson, Oliver, et al.
Pubblicazione: (2026) -
LA-IMR: Latency-Aware, Predictive In-Memory Routing and Proactive Autoscaling for Tail-Latency-Sensitive Cloud Robotics
di: Seo, Eunil, et al.
Pubblicazione: (2025) -
ASA -- The Adaptive Scheduling Algorithm
di: Souza, Abel, et al.
Pubblicazione: (2024)