Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dai, Yinwei, Pan, Rui, Iyer, Anand, Li, Kai, Netravali, Ravi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Aragog: Just-in-Time Model Routing for Scalable Serving of Agentic Workflows
von: Dai, Yinwei, et al.
Veröffentlicht: (2025)
von: Dai, Yinwei, et al.
Veröffentlicht: (2025)
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
Kairos: A Scalable Serving System for Physical AI
von: Dai, Yinwei, et al.
Veröffentlicht: (2026)
von: Dai, Yinwei, et al.
Veröffentlicht: (2026)
Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs
von: Pan, Rui, et al.
Veröffentlicht: (2025)
von: Pan, Rui, et al.
Veröffentlicht: (2025)
OMEGA: A Low-Latency GNN Serving System for Large Graphs
von: Kim, Geon-Woo, et al.
Veröffentlicht: (2025)
von: Kim, Geon-Woo, et al.
Veröffentlicht: (2025)
Taming the Memory Beast: Strategies for Reliable ML Training on Kubernetes
von: Ray, Jaideep
Veröffentlicht: (2024)
von: Ray, Jaideep
Veröffentlicht: (2024)
Dynamic Rebatching for Efficient Early-Exit Inference with DREX
von: Liu, Xuting, et al.
Veröffentlicht: (2025)
von: Liu, Xuting, et al.
Veröffentlicht: (2025)
Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert Offloading
von: Yu, Hanfei, et al.
Veröffentlicht: (2025)
von: Yu, Hanfei, et al.
Veröffentlicht: (2025)
Injecting Adrenaline into LLM Serving: Boosting Resource Utilization and Throughput via Attention Disaggregation
von: Liang, Yunkai, et al.
Veröffentlicht: (2025)
von: Liang, Yunkai, et al.
Veröffentlicht: (2025)
Marconi: Prefix Caching for the Era of Hybrid LLMs
von: Pan, Rui, et al.
Veröffentlicht: (2024)
von: Pan, Rui, et al.
Veröffentlicht: (2024)
OmniInfer: System-Wide Acceleration Techniques for Optimizing LLM Serving Throughput and Latency
von: Wang, Jun, et al.
Veröffentlicht: (2025)
von: Wang, Jun, et al.
Veröffentlicht: (2025)
eMoE: Task-aware Memory Efficient Mixture-of-Experts-Based (MoE) Model Inference
von: Tairin, Suraiya, et al.
Veröffentlicht: (2025)
von: Tairin, Suraiya, et al.
Veröffentlicht: (2025)
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
ReInc: Scaling Training of Dynamic Graph Neural Networks
von: Guan, Mingyu, et al.
Veröffentlicht: (2025)
von: Guan, Mingyu, et al.
Veröffentlicht: (2025)
CascadeServe: Unlocking Model Cascades for Inference Serving
von: Kossmann, Ferdi, et al.
Veröffentlicht: (2024)
von: Kossmann, Ferdi, et al.
Veröffentlicht: (2024)
PolyServe: Efficient Multi-SLO Serving at Scale
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
von: Chen, Siyuan, et al.
Veröffentlicht: (2025)
von: Chen, Siyuan, et al.
Veröffentlicht: (2025)
EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism
von: Chen, Yanxi, et al.
Veröffentlicht: (2023)
von: Chen, Yanxi, et al.
Veröffentlicht: (2023)
Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection
von: Gupta, Vima, et al.
Veröffentlicht: (2024)
von: Gupta, Vima, et al.
Veröffentlicht: (2024)
TetriServe: Efficient DiT Serving for Heterogeneous Image Generation
von: Lu, Runyu, et al.
Veröffentlicht: (2025)
von: Lu, Runyu, et al.
Veröffentlicht: (2025)
Federated Learning for Collaborative Inference Systems: The Case of Early Exit Networks
von: Kaplan, Caelin, et al.
Veröffentlicht: (2024)
von: Kaplan, Caelin, et al.
Veröffentlicht: (2024)
Prompt-Aware Scheduling for Low-Latency LLM Serving
von: Tao, Yiheng, et al.
Veröffentlicht: (2025)
von: Tao, Yiheng, et al.
Veröffentlicht: (2025)
ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache
von: Wang, Shao, et al.
Veröffentlicht: (2026)
von: Wang, Shao, et al.
Veröffentlicht: (2026)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
von: Qiao, Yifan, et al.
Veröffentlicht: (2024)
von: Qiao, Yifan, et al.
Veröffentlicht: (2024)
Distributed Inference on Mobile Edge and Cloud: An Early Exit based Clustering Approach
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024)
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024)
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
von: Wu, Bingyang, et al.
Veröffentlicht: (2024)
von: Wu, Bingyang, et al.
Veröffentlicht: (2024)
FineServe: Precision-Aware KV Slab and Two-Level Scheduling for Heterogeneous Precision LLM Serving
von: Bin, Kyungmin, et al.
Veröffentlicht: (2025)
von: Bin, Kyungmin, et al.
Veröffentlicht: (2025)
vTensor: Flexible Virtual Tensor Management for Efficient LLM Serving
von: Xu, Jiale, et al.
Veröffentlicht: (2024)
von: Xu, Jiale, et al.
Veröffentlicht: (2024)
Magneton: Optimizing Energy Efficiency of ML Systems via Differential Energy Debugging
von: Pan, Yi, et al.
Veröffentlicht: (2025)
von: Pan, Yi, et al.
Veröffentlicht: (2025)
DistrEE: Distributed Early Exit of Deep Neural Network Inference on Edge Devices
von: Peng, Xian, et al.
Veröffentlicht: (2025)
von: Peng, Xian, et al.
Veröffentlicht: (2025)
Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations
von: Li, Shigang, et al.
Veröffentlicht: (2019)
von: Li, Shigang, et al.
Veröffentlicht: (2019)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
von: Su, Zhaoyuan, et al.
Veröffentlicht: (2025)
von: Su, Zhaoyuan, et al.
Veröffentlicht: (2025)
Niyama : Breaking the Silos of LLM Inference Serving
von: Goel, Kanishk, et al.
Veröffentlicht: (2025)
von: Goel, Kanishk, et al.
Veröffentlicht: (2025)
Stateful Large Language Model Serving with Pensieve
von: Yu, Lingfan, et al.
Veröffentlicht: (2023)
von: Yu, Lingfan, et al.
Veröffentlicht: (2023)
Locality-aware Fair Scheduling in LLM Serving
von: Cao, Shiyi, et al.
Veröffentlicht: (2025)
von: Cao, Shiyi, et al.
Veröffentlicht: (2025)
Towards Sustainable Large Language Model Serving
von: Nguyen, Sophia, et al.
Veröffentlicht: (2024)
von: Nguyen, Sophia, et al.
Veröffentlicht: (2024)
Falcon: Advancing Asynchronous BFT Consensus for Lower Latency and Enhanced Throughput
von: Dai, Xiaohai, et al.
Veröffentlicht: (2025)
von: Dai, Xiaohai, et al.
Veröffentlicht: (2025)
HeteroSwitch: Characterizing and Taming System-Induced Data Heterogeneity in Federated Learning
von: Kim, Gyudong, et al.
Veröffentlicht: (2024)
von: Kim, Gyudong, et al.
Veröffentlicht: (2024)
Taming the Titans: A Survey of Efficient LLM Inference Serving
von: Zhen, Ranran, et al.
Veröffentlicht: (2025)
von: Zhen, Ranran, et al.
Veröffentlicht: (2025)
Fast Distributed Inference Serving for Large Language Models
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Aragog: Just-in-Time Model Routing for Scalable Serving of Agentic Workflows
von: Dai, Yinwei, et al.
Veröffentlicht: (2025) -
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
von: Agrawal, Amey, et al.
Veröffentlicht: (2024) -
Kairos: A Scalable Serving System for Physical AI
von: Dai, Yinwei, et al.
Veröffentlicht: (2026) -
Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs
von: Pan, Rui, et al.
Veröffentlicht: (2025) -
OMEGA: A Low-Latency GNN Serving System for Large Graphs
von: Kim, Geon-Woo, et al.
Veröffentlicht: (2025)