Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation
Fuente:
arXiv
Salvato in:
| Autori principali: | Hu, Tiancheng, Qin, Jin, Wang, Zheng, Hu, Junhao, Wang, Yuzheng, Chen, Lei, Shan, Yizhou, Zhang, Mingxing, Cao, Ting, Xia, Chunwei, Cui, Huimin, Xie, Tao, Wang, Chenxi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
di: Hu, Tiancheng, et al.
Pubblicazione: (2026)
di: Hu, Tiancheng, et al.
Pubblicazione: (2026)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
Lotus: Optimizing Disaggregated Transactions with Disaggregated Locks
di: Hu, Zhisheng, et al.
Pubblicazione: (2025)
di: Hu, Zhisheng, et al.
Pubblicazione: (2025)
Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation
di: Chen, Shaoyuan, et al.
Pubblicazione: (2024)
di: Chen, Shaoyuan, et al.
Pubblicazione: (2024)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
di: Chen, Xing, et al.
Pubblicazione: (2025)
di: Chen, Xing, et al.
Pubblicazione: (2025)
Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
di: Liu, Yunzhao, et al.
Pubblicazione: (2025)
di: Liu, Yunzhao, et al.
Pubblicazione: (2025)
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
di: Shi, Tianyao, et al.
Pubblicazione: (2024)
di: Shi, Tianyao, et al.
Pubblicazione: (2024)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
di: Zhang, Zili, et al.
Pubblicazione: (2024)
di: Zhang, Zili, et al.
Pubblicazione: (2024)
EPIC: Efficient Position-Independent Caching for Serving Large Language Models
di: Hu, Junhao, et al.
Pubblicazione: (2024)
di: Hu, Junhao, et al.
Pubblicazione: (2024)
Proceedings of 3rd Workshop on Heterogeneous Composable and Disaggregated Systems
di: Pinto, Christian, et al.
Pubblicazione: (2024)
di: Pinto, Christian, et al.
Pubblicazione: (2024)
Tesserae: Scalable Placement Policies for Deep Learning Workloads
di: Bian, Song, et al.
Pubblicazione: (2025)
di: Bian, Song, et al.
Pubblicazione: (2025)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
di: Ekelund, Jonah, et al.
Pubblicazione: (2025)
di: Ekelund, Jonah, et al.
Pubblicazione: (2025)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
di: Wang, Qipeng
Pubblicazione: (2026)
di: Wang, Qipeng
Pubblicazione: (2026)
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
FlexKV: Flexible Index Offloading for Memory-Disaggregated Key-Value Store
di: Hu, Zhisheng, et al.
Pubblicazione: (2025)
di: Hu, Zhisheng, et al.
Pubblicazione: (2025)
TENT: A Declarative Slice Spraying Engine for Performant and Resilient Data Movement in Disaggregated LLM Serving
di: Ren, Feng, et al.
Pubblicazione: (2026)
di: Ren, Feng, et al.
Pubblicazione: (2026)
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
di: Dwaraknath, Rajat Vadiraj, et al.
Pubblicazione: (2026)
di: Dwaraknath, Rajat Vadiraj, et al.
Pubblicazione: (2026)
Popcorn: Accelerating Kernel K-means on GPUs through Sparse Linear Algebra
di: Bellavita, Julian, et al.
Pubblicazione: (2025)
di: Bellavita, Julian, et al.
Pubblicazione: (2025)
KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
di: Zhang, Huawei, et al.
Pubblicazione: (2025)
di: Zhang, Huawei, et al.
Pubblicazione: (2025)
A Tale of Two Paths: Toward a Hybrid Data Plane for Efficient Far-Memory Applications
di: Chen, Lei, et al.
Pubblicazione: (2024)
di: Chen, Lei, et al.
Pubblicazione: (2024)
FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
di: He, Xuan, et al.
Pubblicazione: (2025)
di: He, Xuan, et al.
Pubblicazione: (2025)
Astra: Efficient and Money-saving Automatic Parallel Strategies Search on Heterogeneous GPUs
di: Wang, Peiran, et al.
Pubblicazione: (2025)
di: Wang, Peiran, et al.
Pubblicazione: (2025)
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
di: Basit, Omar, et al.
Pubblicazione: (2026)
di: Basit, Omar, et al.
Pubblicazione: (2026)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
di: Zhang, Zhexiang, et al.
Pubblicazione: (2025)
di: Zhang, Zhexiang, et al.
Pubblicazione: (2025)
Efficient Multi-round LLM Inference over Disaggregated Serving
di: He, Wenhao, et al.
Pubblicazione: (2026)
di: He, Wenhao, et al.
Pubblicazione: (2026)
RSH-SpMM: A Row-Structured Hybrid Kernel for Sparse Matrix-Matrix Multiplication on GPUs
di: Li, Aiying, et al.
Pubblicazione: (2026)
di: Li, Aiying, et al.
Pubblicazione: (2026)
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
di: Wu, Tianyuan, et al.
Pubblicazione: (2025)
di: Wu, Tianyuan, et al.
Pubblicazione: (2025)
DiFache: Efficient and Scalable Caching on Disaggregated Memory using Decentralized Coherence
di: Zhang, Hanze, et al.
Pubblicazione: (2025)
di: Zhang, Hanze, et al.
Pubblicazione: (2025)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
di: Zhong, Yinmin, et al.
Pubblicazione: (2024)
di: Zhong, Yinmin, et al.
Pubblicazione: (2024)
StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
di: Zhong, Yinmin, et al.
Pubblicazione: (2025)
di: Zhong, Yinmin, et al.
Pubblicazione: (2025)
FastCHGNet: Training one Universal Interatomic Potential to 1.5 Hours with 32 GPUs
di: Zhou, Yuanchang, et al.
Pubblicazione: (2024)
di: Zhou, Yuanchang, et al.
Pubblicazione: (2024)
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
di: Gao, Wei, et al.
Pubblicazione: (2026)
di: Gao, Wei, et al.
Pubblicazione: (2026)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
di: Wang, Chao, et al.
Pubblicazione: (2025)
di: Wang, Chao, et al.
Pubblicazione: (2025)
Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed Clusters
di: Strati, Foteini, et al.
Pubblicazione: (2025)
di: Strati, Foteini, et al.
Pubblicazione: (2025)
vLLM-Omni: Fully Disaggregated Serving for Any-to-Any Multimodal Models
di: Yin, Peiqi, et al.
Pubblicazione: (2026)
di: Yin, Peiqi, et al.
Pubblicazione: (2026)
Efficient Graph-Based Approximate Nearest Neighbor Search Achieving: Low Latency Without Throughput Loss
di: Luo, Jingjia, et al.
Pubblicazione: (2025)
di: Luo, Jingjia, et al.
Pubblicazione: (2025)
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
di: Jiang, Youhe, et al.
Pubblicazione: (2026)
di: Jiang, Youhe, et al.
Pubblicazione: (2026)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
di: Liao, Junhan, et al.
Pubblicazione: (2025)
di: Liao, Junhan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
di: Hu, Cunchen, et al.
Pubblicazione: (2024) -
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
di: Hu, Tiancheng, et al.
Pubblicazione: (2026) -
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
di: Hu, Cunchen, et al.
Pubblicazione: (2024) -
Lotus: Optimizing Disaggregated Transactions with Disaggregated Locks
di: Hu, Zhisheng, et al.
Pubblicazione: (2025) -
Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation
di: Chen, Shaoyuan, et al.
Pubblicazione: (2024)