Proceedings of 3rd Workshop on Heterogeneous Composable and Disaggregated Systems
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Pinto, Christian, Li, Dong, Doudali, Thaleia Dimitra, Giannoula, Christina, Ren, Jie |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
TCM-Serve: Modality-aware Scheduling for Multimodal Large Language Model Inference
par: Papaioannou, Konstantinos, et autres
Publié: (2026)
par: Papaioannou, Konstantinos, et autres
Publié: (2026)
PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems
par: Pennas, Panagiotis Georgios, et autres
Publié: (2026)
par: Pennas, Panagiotis Georgios, et autres
Publié: (2026)
HuntMS: A Framework for Microservice Geo-Distribution for Carbon and Cost Reduction
par: Christofidi, Georgia, et autres
Publié: (2026)
par: Christofidi, Georgia, et autres
Publié: (2026)
DOLMA: A Data Object Level Memory Disaggregation Framework for HPC Applications
par: Zheng, Haoyu, et autres
Publié: (2025)
par: Zheng, Haoyu, et autres
Publié: (2025)
SmartPQ: An Adaptive Concurrent Priority Queue for NUMA Architectures
par: Giannoula, Christina, et autres
Publié: (2024)
par: Giannoula, Christina, et autres
Publié: (2024)
ScalePool: Hybrid XLink-CXL Fabric for Composable Resource Disaggregation in Unified Scale-up Domains
par: Woo, Hyein, et autres
Publié: (2025)
par: Woo, Hyein, et autres
Publié: (2025)
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
par: Patke, Archit, et autres
Publié: (2025)
par: Patke, Archit, et autres
Publié: (2025)
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
par: Jiang, Youhe, et autres
Publié: (2025)
par: Jiang, Youhe, et autres
Publié: (2025)
Lotus: Optimizing Disaggregated Transactions with Disaggregated Locks
par: Hu, Zhisheng, et autres
Publié: (2025)
par: Hu, Zhisheng, et autres
Publié: (2025)
Proceedings of the 1st International Workshop on Low Carbon Computing (LOCO 2024)
par: Vanderbauwhede, Wim, et autres
Publié: (2026)
par: Vanderbauwhede, Wim, et autres
Publié: (2026)
Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
par: Liu, Yunzhao, et autres
Publié: (2025)
par: Liu, Yunzhao, et autres
Publié: (2025)
DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
par: Zhang, Zili, et autres
Publié: (2024)
par: Zhang, Zili, et autres
Publié: (2024)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
par: Wu, Yu, et autres
Publié: (2025)
par: Wu, Yu, et autres
Publié: (2025)
Revealing the Challenges of Attention-FFN Disaggregation for Modern MoE Models and Hardware Systems
par: Liu, Guowei, et autres
Publié: (2026)
par: Liu, Guowei, et autres
Publié: (2026)
Optimising Virtual Resource Mapping in Multi-Level NUMA Disaggregated Systems
par: Lakew, Ewnetu Bayuh, et autres
Publié: (2025)
par: Lakew, Ewnetu Bayuh, et autres
Publié: (2025)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
par: Dong, Xianzhe, et autres
Publié: (2025)
par: Dong, Xianzhe, et autres
Publié: (2025)
Proceedings First Workshop on Adaptable Cloud Architectures
par: De Palma, Giuseppe, et autres
Publié: (2025)
par: De Palma, Giuseppe, et autres
Publié: (2025)
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
par: Li, Rongzhi, et autres
Publié: (2025)
par: Li, Rongzhi, et autres
Publié: (2025)
Software Resource Disaggregation for HPC with Serverless Computing
par: Copik, Marcin, et autres
Publié: (2024)
par: Copik, Marcin, et autres
Publié: (2024)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
par: Ruan, Chaoyi, et autres
Publié: (2025)
par: Ruan, Chaoyi, et autres
Publié: (2025)
Asymmetric Grid Quorum Systems for Heterogeneous Processes
par: Senn, Michael, et autres
Publié: (2025)
par: Senn, Michael, et autres
Publié: (2025)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
par: Chen, Xing, et autres
Publié: (2025)
par: Chen, Xing, et autres
Publié: (2025)
Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation
par: Hu, Tiancheng, et autres
Publié: (2026)
par: Hu, Tiancheng, et autres
Publié: (2026)
Towards Disaggregation-Native Data Streaming between Devices
par: Asmussen, Nils, et autres
Publié: (2024)
par: Asmussen, Nils, et autres
Publié: (2024)
SWARM: Replicating Shared Disaggregated-Memory Data in No Time
par: Murat, Antoine, et autres
Publié: (2024)
par: Murat, Antoine, et autres
Publié: (2024)
DRackSim: Simulator for Rack-scale Memory Disaggregation
par: Puri, Amit, et autres
Publié: (2023)
par: Puri, Amit, et autres
Publié: (2023)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
par: Kumar, Madabattula Rajesh, et autres
Publié: (2025)
par: Kumar, Madabattula Rajesh, et autres
Publié: (2025)
Reconfigurable Heterogeneous Quorum Systems
par: Li, Xiao, et autres
Publié: (2023)
par: Li, Xiao, et autres
Publié: (2023)
Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation
par: Chen, Shaoyuan, et autres
Publié: (2024)
par: Chen, Shaoyuan, et autres
Publié: (2024)
DecLock: A Case of Decoupled Locking for Disaggregated Memory
par: Zhang, Hanze, et autres
Publié: (2025)
par: Zhang, Hanze, et autres
Publié: (2025)
Efficient Multi-round LLM Inference over Disaggregated Serving
par: He, Wenhao, et autres
Publié: (2026)
par: He, Wenhao, et autres
Publié: (2026)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
par: Zhang, Zhexiang, et autres
Publié: (2025)
par: Zhang, Zhexiang, et autres
Publié: (2025)
TENT: A Declarative Slice Spraying Engine for Performant and Resilient Data Movement in Disaggregated LLM Serving
par: Ren, Feng, et autres
Publié: (2026)
par: Ren, Feng, et autres
Publié: (2026)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
par: Chen, Hongyu, et autres
Publié: (2026)
par: Chen, Hongyu, et autres
Publié: (2026)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
par: Hu, Cunchen, et autres
Publié: (2024)
par: Hu, Cunchen, et autres
Publié: (2024)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
par: Wang, Qipeng
Publié: (2026)
par: Wang, Qipeng
Publié: (2026)
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
par: Wu, Tianyuan, et autres
Publié: (2025)
par: Wu, Tianyuan, et autres
Publié: (2025)
PULSE: Accelerating Distributed Pointer-Traversals on Disaggregated Memory (Extended Version)
par: Tang, Yupeng, et autres
Publié: (2023)
par: Tang, Yupeng, et autres
Publié: (2023)
StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
par: Zhong, Yinmin, et autres
Publié: (2025)
par: Zhong, Yinmin, et autres
Publié: (2025)
Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters
par: Chang, Zihan, et autres
Publié: (2024)
par: Chang, Zihan, et autres
Publié: (2024)
Documents similaires
-
TCM-Serve: Modality-aware Scheduling for Multimodal Large Language Model Inference
par: Papaioannou, Konstantinos, et autres
Publié: (2026) -
PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems
par: Pennas, Panagiotis Georgios, et autres
Publié: (2026) -
HuntMS: A Framework for Microservice Geo-Distribution for Carbon and Cost Reduction
par: Christofidi, Georgia, et autres
Publié: (2026) -
DOLMA: A Data Object Level Memory Disaggregation Framework for HPC Applications
par: Zheng, Haoyu, et autres
Publié: (2025) -
SmartPQ: An Adaptive Concurrent Priority Queue for NUMA Architectures
par: Giannoula, Christina, et autres
Publié: (2024)