Strait: Perceiving Priority and Interference in ML Inference Serving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Haidong, Georgantas, Nikolaos |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ML Inference Scheduling with Predictable Latency
von: Zhao, Haidong, et al.
Veröffentlicht: (2025)
von: Zhao, Haidong, et al.
Veröffentlicht: (2025)
Gradient Boosted Risk Scores
von: Georgantas, Costa, et al.
Veröffentlicht: (2026)
von: Georgantas, Costa, et al.
Veröffentlicht: (2026)
Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving
von: Gao, Shihong, et al.
Veröffentlicht: (2025)
von: Gao, Shihong, et al.
Veröffentlicht: (2025)
CascadeServe: Unlocking Model Cascades for Inference Serving
von: Kossmann, Ferdi, et al.
Veröffentlicht: (2024)
von: Kossmann, Ferdi, et al.
Veröffentlicht: (2024)
Priority-Aware Preemptive Scheduling for Mixed-Priority Workloads in MoE Inference
von: Siavashi, Mohammad, et al.
Veröffentlicht: (2025)
von: Siavashi, Mohammad, et al.
Veröffentlicht: (2025)
T-TAMER: Provably Taming Trade-offs in ML Serving
von: Yang, Yuanyuan, et al.
Veröffentlicht: (2025)
von: Yang, Yuanyuan, et al.
Veröffentlicht: (2025)
Collaborative Speculative Inference for Efficient LLM Inference Serving
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching
von: Zhao, Yilong, et al.
Veröffentlicht: (2024)
von: Zhao, Yilong, et al.
Veröffentlicht: (2024)
Toward Robust and Efficient ML-Based GPU Caching for Modern Inference
von: Chen, Peng, et al.
Veröffentlicht: (2025)
von: Chen, Peng, et al.
Veröffentlicht: (2025)
FluidML: Fast and Memory Efficient Inference Optimization
von: Liu, Jinjie, et al.
Veröffentlicht: (2024)
von: Liu, Jinjie, et al.
Veröffentlicht: (2024)
ShadowServe: Interference-Free KV Cache Fetching for Distributed Prefix Caching
von: Xiang, Xingyu, et al.
Veröffentlicht: (2025)
von: Xiang, Xingyu, et al.
Veröffentlicht: (2025)
Priority-Aware Model-Distributed Inference at Edge Networks
von: Li, Teng, et al.
Veröffentlicht: (2024)
von: Li, Teng, et al.
Veröffentlicht: (2024)
Spatial Deconfounder: Interference-Aware Deconfounding for Spatial Causal Inference
von: Khot, Ayush, et al.
Veröffentlicht: (2025)
von: Khot, Ayush, et al.
Veröffentlicht: (2025)
Integrating Active Learning in Causal Inference with Interference: A Novel Approach in Online Experiments
von: Zhu, Hongtao, et al.
Veröffentlicht: (2024)
von: Zhu, Hongtao, et al.
Veröffentlicht: (2024)
Accelerating TinyML Inference on Microcontrollers through Approximate Kernels
von: Armeniakos, Giorgos, et al.
Veröffentlicht: (2024)
von: Armeniakos, Giorgos, et al.
Veröffentlicht: (2024)
SCORPIO: Serving the Right Requests at the Right Time for Heterogeneous SLOs in LLM Inference
von: Tang, Yinghao, et al.
Veröffentlicht: (2025)
von: Tang, Yinghao, et al.
Veröffentlicht: (2025)
Applied Causal Inference Powered by ML and AI
von: Chernozhukov, Victor, et al.
Veröffentlicht: (2024)
von: Chernozhukov, Victor, et al.
Veröffentlicht: (2024)
Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving
von: Dai, Yinwei, et al.
Veröffentlicht: (2023)
von: Dai, Yinwei, et al.
Veröffentlicht: (2023)
Reasoning Language Model Inference Serving Unveiled: An Empirical Study
von: Li, Qi, et al.
Veröffentlicht: (2025)
von: Li, Qi, et al.
Veröffentlicht: (2025)
Neighborhood Adaptive Estimators for Causal Inference under Network Interference
von: Belloni, Alexandre, et al.
Veröffentlicht: (2022)
von: Belloni, Alexandre, et al.
Veröffentlicht: (2022)
Biathlon: Harnessing Model Resilience for Accelerating ML Inference Pipelines
von: Chang, Chaokun, et al.
Veröffentlicht: (2024)
von: Chang, Chaokun, et al.
Veröffentlicht: (2024)
Scalable AI Inference: Performance Analysis and Optimization of AI Model Serving
von: Pham, Hung Cuong, et al.
Veröffentlicht: (2026)
von: Pham, Hung Cuong, et al.
Veröffentlicht: (2026)
Fast Distributed Inference Serving for Large Language Models
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
Computing Within Limits: An Empirical Study of Energy Consumption in ML Training and Inference
von: Mavromatis, Ioannis, et al.
Veröffentlicht: (2024)
von: Mavromatis, Ioannis, et al.
Veröffentlicht: (2024)
Double Machine Learning for Causal Inference under Shared-State Interference
von: Hays, Chris, et al.
Veröffentlicht: (2025)
von: Hays, Chris, et al.
Veröffentlicht: (2025)
Priority-Aware Shapley Value
von: Lee, Kiljae, et al.
Veröffentlicht: (2026)
von: Lee, Kiljae, et al.
Veröffentlicht: (2026)
InferF: Declarative Factorization of AI/ML Inferences over Joins
von: Chowdhury, Kanchan, et al.
Veröffentlicht: (2025)
von: Chowdhury, Kanchan, et al.
Veröffentlicht: (2025)
Niyama : Breaking the Silos of LLM Inference Serving
von: Goel, Kanishk, et al.
Veröffentlicht: (2025)
von: Goel, Kanishk, et al.
Veröffentlicht: (2025)
On Evaluating Performance of LLM Inference Serving Systems
von: Agrawal, Amey, et al.
Veröffentlicht: (2025)
von: Agrawal, Amey, et al.
Veröffentlicht: (2025)
Practical and Private Hybrid ML Inference with Fully Homomorphic Encryption
von: Biswas, Sayan, et al.
Veröffentlicht: (2025)
von: Biswas, Sayan, et al.
Veröffentlicht: (2025)
HELIOS: Adaptive Model And Early-Exit Selection for Efficient LLM Inference Serving
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
von: Ziller, Thomas, et al.
Veröffentlicht: (2026)
von: Ziller, Thomas, et al.
Veröffentlicht: (2026)
Design-Based Bandits Under Network Interference: Trade-Off Between Regret and Statistical Inference
von: Wang, Zichen, et al.
Veröffentlicht: (2025)
von: Wang, Zichen, et al.
Veröffentlicht: (2025)
The CAP Principle for LLM Serving: A Survey of Long-Context Large Language Model Serving
von: Zeng, Pai, et al.
Veröffentlicht: (2024)
von: Zeng, Pai, et al.
Veröffentlicht: (2024)
The Local Approach to Causal Inference under Network Interference
von: Auerbach, Eric, et al.
Veröffentlicht: (2021)
von: Auerbach, Eric, et al.
Veröffentlicht: (2021)
MicroFlow: An Efficient Rust-Based Inference Engine for TinyML
von: Carnelos, Matteo, et al.
Veröffentlicht: (2024)
von: Carnelos, Matteo, et al.
Veröffentlicht: (2024)
The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization
von: Chung, Jae-Won, et al.
Veröffentlicht: (2025)
von: Chung, Jae-Won, et al.
Veröffentlicht: (2025)
Federated Inference: Toward Privacy-Preserving Collaborative and Incentivized Model Serving
von: Seo, Jungwon, et al.
Veröffentlicht: (2026)
von: Seo, Jungwon, et al.
Veröffentlicht: (2026)
Obsidian: Cooperative State-Space Exploration for Performant Inference on Secure ML Accelerators
von: Banerjee, Sarbartha, et al.
Veröffentlicht: (2024)
von: Banerjee, Sarbartha, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ML Inference Scheduling with Predictable Latency
von: Zhao, Haidong, et al.
Veröffentlicht: (2025) -
Gradient Boosted Risk Scores
von: Georgantas, Costa, et al.
Veröffentlicht: (2026) -
Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving
von: Gao, Shihong, et al.
Veröffentlicht: (2025) -
CascadeServe: Unlocking Model Cascades for Inference Serving
von: Kossmann, Ferdi, et al.
Veröffentlicht: (2024) -
Priority-Aware Preemptive Scheduling for Mixed-Priority Workloads in MoE Inference
von: Siavashi, Mohammad, et al.
Veröffentlicht: (2025)