Strait: Perceiving Priority and Interference in ML Inference Serving
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Haidong, Georgantas, Nikolaos |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ML Inference Scheduling with Predictable Latency
by: Zhao, Haidong, et al.
Published: (2025)
by: Zhao, Haidong, et al.
Published: (2025)
Gradient Boosted Risk Scores
by: Georgantas, Costa, et al.
Published: (2026)
by: Georgantas, Costa, et al.
Published: (2026)
Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving
by: Gao, Shihong, et al.
Published: (2025)
by: Gao, Shihong, et al.
Published: (2025)
CascadeServe: Unlocking Model Cascades for Inference Serving
by: Kossmann, Ferdi, et al.
Published: (2024)
by: Kossmann, Ferdi, et al.
Published: (2024)
Priority-Aware Preemptive Scheduling for Mixed-Priority Workloads in MoE Inference
by: Siavashi, Mohammad, et al.
Published: (2025)
by: Siavashi, Mohammad, et al.
Published: (2025)
T-TAMER: Provably Taming Trade-offs in ML Serving
by: Yang, Yuanyuan, et al.
Published: (2025)
by: Yang, Yuanyuan, et al.
Published: (2025)
Collaborative Speculative Inference for Efficient LLM Inference Serving
by: Gao, Luyao, et al.
Published: (2025)
by: Gao, Luyao, et al.
Published: (2025)
BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching
by: Zhao, Yilong, et al.
Published: (2024)
by: Zhao, Yilong, et al.
Published: (2024)
Toward Robust and Efficient ML-Based GPU Caching for Modern Inference
by: Chen, Peng, et al.
Published: (2025)
by: Chen, Peng, et al.
Published: (2025)
FluidML: Fast and Memory Efficient Inference Optimization
by: Liu, Jinjie, et al.
Published: (2024)
by: Liu, Jinjie, et al.
Published: (2024)
ShadowServe: Interference-Free KV Cache Fetching for Distributed Prefix Caching
by: Xiang, Xingyu, et al.
Published: (2025)
by: Xiang, Xingyu, et al.
Published: (2025)
Priority-Aware Model-Distributed Inference at Edge Networks
by: Li, Teng, et al.
Published: (2024)
by: Li, Teng, et al.
Published: (2024)
Spatial Deconfounder: Interference-Aware Deconfounding for Spatial Causal Inference
by: Khot, Ayush, et al.
Published: (2025)
by: Khot, Ayush, et al.
Published: (2025)
Integrating Active Learning in Causal Inference with Interference: A Novel Approach in Online Experiments
by: Zhu, Hongtao, et al.
Published: (2024)
by: Zhu, Hongtao, et al.
Published: (2024)
Accelerating TinyML Inference on Microcontrollers through Approximate Kernels
by: Armeniakos, Giorgos, et al.
Published: (2024)
by: Armeniakos, Giorgos, et al.
Published: (2024)
SCORPIO: Serving the Right Requests at the Right Time for Heterogeneous SLOs in LLM Inference
by: Tang, Yinghao, et al.
Published: (2025)
by: Tang, Yinghao, et al.
Published: (2025)
Applied Causal Inference Powered by ML and AI
by: Chernozhukov, Victor, et al.
Published: (2024)
by: Chernozhukov, Victor, et al.
Published: (2024)
Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving
by: Dai, Yinwei, et al.
Published: (2023)
by: Dai, Yinwei, et al.
Published: (2023)
Reasoning Language Model Inference Serving Unveiled: An Empirical Study
by: Li, Qi, et al.
Published: (2025)
by: Li, Qi, et al.
Published: (2025)
Neighborhood Adaptive Estimators for Causal Inference under Network Interference
by: Belloni, Alexandre, et al.
Published: (2022)
by: Belloni, Alexandre, et al.
Published: (2022)
Biathlon: Harnessing Model Resilience for Accelerating ML Inference Pipelines
by: Chang, Chaokun, et al.
Published: (2024)
by: Chang, Chaokun, et al.
Published: (2024)
Scalable AI Inference: Performance Analysis and Optimization of AI Model Serving
by: Pham, Hung Cuong, et al.
Published: (2026)
by: Pham, Hung Cuong, et al.
Published: (2026)
Fast Distributed Inference Serving for Large Language Models
by: Wu, Bingyang, et al.
Published: (2023)
by: Wu, Bingyang, et al.
Published: (2023)
Computing Within Limits: An Empirical Study of Energy Consumption in ML Training and Inference
by: Mavromatis, Ioannis, et al.
Published: (2024)
by: Mavromatis, Ioannis, et al.
Published: (2024)
Double Machine Learning for Causal Inference under Shared-State Interference
by: Hays, Chris, et al.
Published: (2025)
by: Hays, Chris, et al.
Published: (2025)
Priority-Aware Shapley Value
by: Lee, Kiljae, et al.
Published: (2026)
by: Lee, Kiljae, et al.
Published: (2026)
InferF: Declarative Factorization of AI/ML Inferences over Joins
by: Chowdhury, Kanchan, et al.
Published: (2025)
by: Chowdhury, Kanchan, et al.
Published: (2025)
Niyama : Breaking the Silos of LLM Inference Serving
by: Goel, Kanishk, et al.
Published: (2025)
by: Goel, Kanishk, et al.
Published: (2025)
On Evaluating Performance of LLM Inference Serving Systems
by: Agrawal, Amey, et al.
Published: (2025)
by: Agrawal, Amey, et al.
Published: (2025)
Practical and Private Hybrid ML Inference with Fully Homomorphic Encryption
by: Biswas, Sayan, et al.
Published: (2025)
by: Biswas, Sayan, et al.
Published: (2025)
HELIOS: Adaptive Model And Early-Exit Selection for Efficient LLM Inference Serving
by: Kumar, Avinash, et al.
Published: (2025)
by: Kumar, Avinash, et al.
Published: (2025)
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
by: Agrawal, Amey, et al.
Published: (2024)
by: Agrawal, Amey, et al.
Published: (2024)
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
by: Ziller, Thomas, et al.
Published: (2026)
by: Ziller, Thomas, et al.
Published: (2026)
The CAP Principle for LLM Serving: A Survey of Long-Context Large Language Model Serving
by: Zeng, Pai, et al.
Published: (2024)
by: Zeng, Pai, et al.
Published: (2024)
Design-Based Bandits Under Network Interference: Trade-Off Between Regret and Statistical Inference
by: Wang, Zichen, et al.
Published: (2025)
by: Wang, Zichen, et al.
Published: (2025)
The Local Approach to Causal Inference under Network Interference
by: Auerbach, Eric, et al.
Published: (2021)
by: Auerbach, Eric, et al.
Published: (2021)
MicroFlow: An Efficient Rust-Based Inference Engine for TinyML
by: Carnelos, Matteo, et al.
Published: (2024)
by: Carnelos, Matteo, et al.
Published: (2024)
The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization
by: Chung, Jae-Won, et al.
Published: (2025)
by: Chung, Jae-Won, et al.
Published: (2025)
Federated Inference: Toward Privacy-Preserving Collaborative and Incentivized Model Serving
by: Seo, Jungwon, et al.
Published: (2026)
by: Seo, Jungwon, et al.
Published: (2026)
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing
by: Gao, Lei, et al.
Published: (2025)
by: Gao, Lei, et al.
Published: (2025)
Similar Items
-
ML Inference Scheduling with Predictable Latency
by: Zhao, Haidong, et al.
Published: (2025) -
Gradient Boosted Risk Scores
by: Georgantas, Costa, et al.
Published: (2026) -
Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving
by: Gao, Shihong, et al.
Published: (2025) -
CascadeServe: Unlocking Model Cascades for Inference Serving
by: Kossmann, Ferdi, et al.
Published: (2024) -
Priority-Aware Preemptive Scheduling for Mixed-Priority Workloads in MoE Inference
by: Siavashi, Mohammad, et al.
Published: (2025)